← Back to the search

How the detective works

An AI flood detective that refuses to guess.

Each pin is a written flood report, not a storm. Many reports were filed by the NWS at a city placeholder. The detective reads each report’s own words and re-files it at the street it names, or says it doesn’t know.

What the AI does here, and what each job has measured

  • The flood detectiveMeasured

    Reads each National Weather Service flood report and re-files a misplaced pin at the street the words name, or says it doesn’t know.

    In a pre-registered test on 100 fresh reports, the AI detective cut wrong pins from 28% to 2%. It places fewer reports (16% vs 33% within 1 km), because it refuses to guess when a report names only a road.

    Real run on this site’s data: it re-filed 131 of 494 imprecise flood reports around the sample and featured addresses, in 8 states, and said “I don’t know” on the other 363. These re-filings are not scored against the truth yet.

    It said “I don’t know” on 79% of reports instead of guessing, and no wrong pin was a confident one. Pilot: 36% → 8%. Real placements are in the data now.

  • The damage mapperPilot result

    Reads a tornado’s official damage narrative and pins the named streets and landmarks where it says there was damage.

    On 5 hand-labelled NWS surveys: extraction precision 0.98, recall 0.86; 45 of 72 damage places pinned.

    68 damage places are on the map now, each with the NWS’s own words. A spot-check of 10 random pins found 10 on the named place and the tornado’s track. A street pin is where that street meets the track, not a spot the survey states. Five surveys is a small test, not a pre-registered one.

  • The briefing and “Ask your street”Measured

    Writes the next-72-hours brief and answers questions, only in sentences a source backs up. Code checks every citation and number.

    98.5% of the sentences it wrote passed the checks (66 of 67, across 20 questions on 4 addresses).

    The rest were removed before anyone saw them. Measured 2026-10-08 with claude-sonnet-5-5. It shows the checks work; it does not show every sentence is wise.

  • The flood forecaster we trained and set asideMeasured

    A small model we trained to predict where flood reports would follow, to see whether we could beat the official outlook.

    0.0% better than the official outlook using only what is known in advance (+3.5% if it is told the rain after the fact).

    So we explain the official forecasts instead of competing with them.

How the detective decides

  1. Read the words

    The NWS report text, not a news story.

  2. Quote the places

    Every place it proposes must appear in the text. It cannot add a cross street.

  3. Check the map

    Only a Census address or intersection match is accepted.

  4. Say “I don’t know”

    Vague or region-wide text stays unplaced.

Pre-registered test, 100 fresh reports Measured

In a pre-registered test on 100 fresh reports, the AI detective cut wrong pins from 28% to 2%.

It places fewer reports (16% vs 33% within 1 km), because it refuses to guess when a report names only a road.

Wrong pins, more than 5 km off (lower is better)

Without AI28%
AI detective2%

Placed within 1 km of the true spot (higher is better)

Without AI33%
AI detective16%

Pre-registered: the sample, the measures and the pass bar were committed to the repository before the run, then it ran once on 100 NWS reports whose true locations were hidden. Wrong pins is the main measure; within 1 km is the reach it gives up. The AI detective said “I don’t know” on 79% of reports (the no-AI baseline left 10% unplaced). Of the 26 reports where exactly one of them was wrong, the AI detective was the wrong one in 0 (McNemar p = 3e-8). No wrong pin was a confident one. The test’s own verdict: fewer wrong pins, less reach.

Pilot: 36% → 8% (an earlier pilot, not pre-registered). Test files: pipeline/eval/score.json, pipeline/eval/PREREG.md.

We trained a flood forecaster. Here is what it taught us.

We trained a small model on 2016 to 2022 and tested it on 2024 to 2025: did a flood get reported within 10 km of 78 city points in 7 states. Told the rain after the fact, it beat the official outlook’s score by 3.5%. But tomorrow’s rain is not known today. Using only what is known in advance, it scored 0.0%: exactly as good as the official outlook, and no better.

Brier skill against the official flood outlook alone, test years 2024 to 2025Zero means no better than the official outlook. Knowing the rain after the fact: plus 3.5 percent, 95 percent range plus 1.7 percent to plus 5.3 percent. Knowing only what is known in advance: 0.0 percent, 95 percent range minus 1.4 percent to plus 1.3 percent.-2%0%+2%+4%+6%Knowing the rain after the fact+3.5%Knowing only what is known in advance0.0%0 = no better than the official outlook
Brier skill against the official outlook alone (higher is better), with its 95% range. Hollow dot: could not run live, because tomorrow’s rain is not known today. Filled dot: uses only what is known in advance.
Text version of this chart
Brier skill against the official outlook alone, tested on 2024 to 2025
ModelSkill95% rangeCould run live
Knowing the rain after the fact+3.5%+1.7% to +5.3%No
Knowing only what is known in advance0.0%-1.4% to +1.3%Yes

So Climate Rewind explains the official forecasts instead of competing with them. That is also what the published research found: Colorado State University’s machine-learning system was less skillful than the human outlook on Day 1.

The details and the limits
Trained, tested
Trained on 2016 to 2022 (199,368 point-days, 920 with a flood report), tested on 2024 to 2025 (56,238 point-days, 350 with a flood report on 206 separate dates).
What “skill” means
Brier skill compares how close the model’s probabilities were to what happened with the official outlook’s. 0% is no better, higher is better. The 95% range comes from resampling whole dates, because one storm makes many reports.
The official outlook itself
On the same test the official outlook beat the usual local rate for the time of year by 6.0% (range 4.5% to 7.6%). That skill is the Weather Service’s.
If the rain is wrong
In a simulation (not real forecast rain), the hindsight advantage shrank from 3.5% to 0.7% when the rain input was moderately wrong, and to about nothing when it was badly wrong.
Limits
The labels are flood reports, which follow where people are and report, not flooding itself. The points are cities in seven states. The test has 350 reporting days. Absolute probabilities drift and would need re-fitting every year. We did not ship it.
Published research
Schumacher et al., Bulletin of the American Meteorological Society, 2021 (doi 10.1175/BAMS-D-20-0186.1): the Colorado State University machine-learning system was less skillful than the operational outlook on Day 1.
Full write-up
docs/forecast-experiment.md in the project repository: how it was set up, every table, and the caveats.
What is real, and what is sample
Radar
NEXRAD composite, 5-minute frames (Iowa Environmental Mesonet)
Map
Esri World Dark Gray tiles (© Esri, HERE, Garmin, © OpenStreetMap contributors). The optional 3D view adds OpenFreeMap (© OpenMapTiles, data from OpenStreetMap) and Mapzen terrain tiles (AWS Terrain Tiles).
Flood reports
NWS Local Storm Reports and NCEI Storm Events, quoted from the archive
Heat
Weather reanalysis (ERA5 via Open-Meteo), daily feels-like highs. A model fitted to observations, not a thermometer at your door. If Open-Meteo is rate limited, the page uses NASA POWER (MERRA-2) instead, computes a heat index from it, and says so where the number appears.
Tornadoes
NOAA Storm Events, including the NWS damage narratives (quoted word for word)
Live strip
api.weather.gov alerts for this address, fetched on load and every 5 minutes. Never faked.
FEMA zone
NFHL point query for this address
AI detective
Where it has not run for an area yet, the page says so and moves no pin. Where it has, a placement tagged Verified is a Census address or intersection match for the words, and Area is a named area’s centre. None of these has been scored against the truth yet.
Sample
The earlier pilot numbers (not the pre-registered result), and anything else that is not real data