How the detective works
An AI flood detective that refuses to guess.
Each pin is a written flood report, not a storm. Many reports were filed by the NWS at a city placeholder. The detective reads each report’s own words and re-files it at the street it names, or says it doesn’t know.
What the AI does here, and what each job has measured
- The flood detectiveMeasured
Reads each National Weather Service flood report and re-files a misplaced pin at the street the words name, or says it doesn’t know.
In a pre-registered test on 100 fresh reports, the AI detective cut wrong pins from 28% to 2%. It places fewer reports (16% vs 33% within 1 km), because it refuses to guess when a report names only a road.
Real run on this site’s data: it re-filed 131 of 494 imprecise flood reports around the sample and featured addresses, in 8 states, and said “I don’t know” on the other 363. These re-filings are not scored against the truth yet.
It said “I don’t know” on 79% of reports instead of guessing, and no wrong pin was a confident one. Pilot: 36% → 8%. Real placements are in the data now.
- The damage mapperPilot result
Reads a tornado’s official damage narrative and pins the named streets and landmarks where it says there was damage.
On 5 hand-labelled NWS surveys: extraction precision 0.98, recall 0.86; 45 of 72 damage places pinned.
68 damage places are on the map now, each with the NWS’s own words. A spot-check of 10 random pins found 10 on the named place and the tornado’s track. A street pin is where that street meets the track, not a spot the survey states. Five surveys is a small test, not a pre-registered one.
- The briefing and “Ask your street”Measured
Writes the next-72-hours brief and answers questions, only in sentences a source backs up. Code checks every citation and number.
98.5% of the sentences it wrote passed the checks (66 of 67, across 20 questions on 4 addresses).
The rest were removed before anyone saw them. Measured 2026-10-08 with claude-sonnet-5-5. It shows the checks work; it does not show every sentence is wise.
- The flood forecaster we trained and set asideMeasured
A small model we trained to predict where flood reports would follow, to see whether we could beat the official outlook.
0.0% better than the official outlook using only what is known in advance (+3.5% if it is told the rain after the fact).
So we explain the official forecasts instead of competing with them.
How the detective decides
- Read the words
The NWS report text, not a news story.
- Quote the places
Every place it proposes must appear in the text. It cannot add a cross street.
- Check the map
Only a Census address or intersection match is accepted.
- Say “I don’t know”
Vague or region-wide text stays unplaced.
Pre-registered test, 100 fresh reports Measured
Pre-registered: the sample, the measures and the pass bar were committed to the repository before the run, then it ran once on 100 NWS reports whose true locations were hidden. Wrong pins is the main measure; within 1 km is the reach it gives up. The AI detective said “I don’t know” on 79% of reports (the no-AI baseline left 10% unplaced). Of the 26 reports where exactly one of them was wrong, the AI detective was the wrong one in 0 (McNemar p = 3e-8). No wrong pin was a confident one. The test’s own verdict: fewer wrong pins, less reach.
Pilot: 36% → 8% (an earlier pilot, not pre-registered). Test files: pipeline/eval/score.json, pipeline/eval/PREREG.md.
We trained a flood forecaster. Here is what it taught us.
We trained a small model on 2016 to 2022 and tested it on 2024 to 2025: did a flood get reported within 10 km of 78 city points in 7 states. Told the rain after the fact, it beat the official outlook’s score by 3.5%. But tomorrow’s rain is not known today. Using only what is known in advance, it scored 0.0%: exactly as good as the official outlook, and no better.
Text version of this chart
| Model | Skill | 95% range | Could run live |
|---|---|---|---|
| Knowing the rain after the fact | +3.5% | +1.7% to +5.3% | No |
| Knowing only what is known in advance | 0.0% | -1.4% to +1.3% | Yes |
So Climate Rewind explains the official forecasts instead of competing with them. That is also what the published research found: Colorado State University’s machine-learning system was less skillful than the human outlook on Day 1.
The details and the limits
- Trained, tested
- Trained on 2016 to 2022 (199,368 point-days, 920 with a flood report), tested on 2024 to 2025 (56,238 point-days, 350 with a flood report on 206 separate dates).
- What “skill” means
- Brier skill compares how close the model’s probabilities were to what happened with the official outlook’s. 0% is no better, higher is better. The 95% range comes from resampling whole dates, because one storm makes many reports.
- The official outlook itself
- On the same test the official outlook beat the usual local rate for the time of year by 6.0% (range 4.5% to 7.6%). That skill is the Weather Service’s.
- If the rain is wrong
- In a simulation (not real forecast rain), the hindsight advantage shrank from 3.5% to 0.7% when the rain input was moderately wrong, and to about nothing when it was badly wrong.
- Limits
- The labels are flood reports, which follow where people are and report, not flooding itself. The points are cities in seven states. The test has 350 reporting days. Absolute probabilities drift and would need re-fitting every year. We did not ship it.
- Published research
- Schumacher et al., Bulletin of the American Meteorological Society, 2021 (doi 10.1175/BAMS-D-20-0186.1): the Colorado State University machine-learning system was less skillful than the operational outlook on Day 1.
- Full write-up
docs/forecast-experiment.mdin the project repository: how it was set up, every table, and the caveats.
What is real, and what is sample
- Radar
- NEXRAD composite, 5-minute frames (Iowa Environmental Mesonet)
- Map
- Esri World Dark Gray tiles (© Esri, HERE, Garmin, © OpenStreetMap contributors). The optional 3D view adds OpenFreeMap (© OpenMapTiles, data from OpenStreetMap) and Mapzen terrain tiles (AWS Terrain Tiles).
- Flood reports
- NWS Local Storm Reports and NCEI Storm Events, quoted from the archive
- Heat
- Weather reanalysis (ERA5 via Open-Meteo), daily feels-like highs. A model fitted to observations, not a thermometer at your door. If Open-Meteo is rate limited, the page uses NASA POWER (MERRA-2) instead, computes a heat index from it, and says so where the number appears.
- Tornadoes
- NOAA Storm Events, including the NWS damage narratives (quoted word for word)
- Live strip
- api.weather.gov alerts for this address, fetched on load and every 5 minutes. Never faked.
- FEMA zone
- NFHL point query for this address
- AI detective
- Where it has not run for an area yet, the page says so and moves no pin. Where it has, a placement tagged Verified is a Census address or intersection match for the words, and Area is a named area’s centre. None of these has been scored against the truth yet.
- Sample
- The earlier pilot numbers (not the pre-registered result), and anything else that is not real data