Would public data have caught it before regulators did?
Hazium builds a temporally-aware knowledge graph over EU pesticide approvals, hazard classifications, and scientific literature, then asks a falsifiable question of it: ranking substances for future regulatory risk using only evidence that existed at the time, measured against real EU bans that happened years later.
Why this exists
Hazium began with a simple question: could publicly available data have revealed a Swedish pesticide controversy before it became national news?
The answer was no, and the reason is precise. Fluazinam's actual concern is groundwater: it breaks down into the PFAS substance trifluoroacetic acid (TFA), which spreads to groundwater. Kemikalieinspektionen opened a formal reevaluation in November 2025, and an SVT investigation made it national news in July 2026. The sources ingested so far, EU regulatory history, EU hazard classifications, and Swedish sales, do not cover groundwater or residue monitoring, so that specific signal sits outside the current data.
The concern itself is not in doubt. A national SGU groundwater investigation across 2023 to 2025 found TFA at 91 percent of 237 sites (median 230 ng/l), tied to fluorinated plant-protection products breaking down. Sweden's historical pesticide monitoring meanwhile records fluazinam at zero of 139 groundwater analyses: the parent never arrives because it becomes TFA. That monitoring is too recent to have fed a pre-2023 ranking, so it is not a model input; it is independent, after-the-fact confirmation of the concern the project set out to anticipate. Folding groundwater and residue monitoring in as a present-day signal is the next step on the roadmap.
That gap is the real origin of the project. Environmental and public health evidence exists in volume across Europe: regulatory decisions, hazard classifications, sales statistics, scientific literature. It is split across agencies that do not share a schema, a timeline, or even a common substance identifier. Hazium joins that evidence into one temporally dated graph, so a ranking can be checked against what was knowable at a real cutoff.
How it decides
Every ranking traces back to real, dated, publicly-sourced facts. A gradient-boosted model (XGBoost) is trained on six feature groups, each grounded in a specific public source:
Hazard classification history
How many severe hazard codes a substance carries under EU CLP: carcinogenicity, aquatic toxicity, reproductive toxicity, and how recently a classification was added.
Scientific assessment scrutiny
How many EFSA toxicological assessments exist, over what span of years. Sustained scientific attention is itself a signal, independent of the conclusion.
Sales and usage trends
Tonnage sold over time, trend direction, and volatility. A substance quietly losing market share behaves differently from one still expanding.
EU regulatory history
How long a substance has held EU approval, and its history of renewals or restrictions: the single strongest signal the model has found so far.
Graph structure
Shared hazard classifications and metabolic degradation links to other substances already flagged as concerning.
Independent literature signal
How a substance's share of hazard-flavoured scientific literature (Europe PMC) compares to the rest of the field in the same year. This is the one signal here that sits upstream of the regulatory process itself.
The model is always compared against trivial baselines: severe-hazard count alone, latest sales tonnage alone, assessment count alone, on the identical task and split. If it doesn't beat them, the baseline becomes the published result.
The result: HEWB
The Hazium Early Warning Benchmark fixes ten historical EU pesticide bans, real regulatory actions, not hypothetical cases. At each annual cutoff from 2009, using only evidence dated before that cutoff, it asks where Hazium would have ranked the substance among every substance the graph knew about that year, roughly 5,900 of them.
Months before the ban is the easy number. The harder question, and the one that shows capability, is whether Hazium was ahead of the independent world: the regulator's first public concern, which arrives long before the final paperwork. The literature signal became a model input, so it is left out of this comparison; what remains are dated regulatory milestones the model never sees.
Click a substance for its use and full rank history.
The landmarks it misses
On the developmental-neurotoxicity and reprotoxic cases, chlorpyrifos, its methyl sister, thiacloprid, and mancozeb, Hazium ranked the substance among the riskiest roughly a decade before EFSA's first public concern. On the neonicotinoids it was early relative to the 2013 EU restriction, though national bans were already emerging. On dimethoate it moved level with the regulator, and on imidacloprid it flagged late; both are on the chart. Epoxiconazole it never flagged at all. Where a substance had a real public controversy, the chart marks that too: Hazium flagged chlorpyrifos years before its 2015 US ban fight, and the neonicotinoids before the 2012 bee campaign. Most landmarks had no public profile at all when Hazium flagged them.
HEWB v1.4. Flag dates come from the frozen benchmark run under strict pre-cutoff evidence discipline; out-of-fold scores are averaged over repeated cross-validation, so the ranks hold steady across resampling. Regulatory milestone dates are hand-verified against the enacting act or EFSA output.
What was knowable, year by year
Every fact carries the date it became public, so this is the evidence available at each cutoff, not what is known today. Substances appear once they are connected to Clothianidin within two steps, most often by sharing a hazard classification.
Click any mark to see what it is. EFSA assessments link to their published opinion.
Principles
Temporal integrity
Every fact and edge carries the earliest date it was publicly knowable, and evaluation sees only facts dated before the cutoff being tested. A claim like “would have flagged it” is measured against what the model could actually have known at the time.
The baseline rule
No learned model is reported without a trivial baseline on the identical task and split. When the baseline wins, it becomes the published result.
Honesty over novelty
HEWB publishes the misses next to the hits. Every version records which landmarks it fails to flag before their real regulatory action.
Evidence paths
A ranking is more than a number. Every score traces through the graph to the documents behind it: an EFSA opinion, an EU regulation, a hazard classification, each one a reader can open and check.