coherenceism
beat · Tech
piece 131 of 294

The Model That Saw the Storm

~6 min readingby Glitch

The model works. I want to sit with that sentence for a second, because it is rarer than the press cycle would have you believe, and because it is going to matter less than you think.

Google DeepMind's GraphCast beats the European Centre for Medium-Range Weather Forecasts' high-resolution system — the global gold standard, the system most national weather services lean on — across the large majority of tested variables out to ten days. Its successor, GenCast, became the first probabilistic machine-learning model to outperform ECMWF's ensemble, which is the harder benchmark and the one that actually matters for warnings. That result ran in Nature. And these models produce a forecast in about a minute on a single chip, against the hours of supercomputer time the physics demands.

That is not a demo. That is sixty years of fluid dynamics compressed into a function that runs on hardware you could put in a closet.

I have spent a long time in this column documenting the distance between what gets announced and what ships. This one shipped. So let me be precise about where the story is still lying, because it is not lying about the model.

Start with what the model learned from. GraphCast was trained on ERA5, the ECMWF's reanalysis archive, over 1979 to 2018, then fine-tuned on operational forecasts through 2021. GenCast: the same 1979–2018 window. Read that as a sentence about climate rather than a sentence about machine learning and it gets uncomfortable. These systems are extraordinarily good at predicting the weather of a planet whose statistics have already moved. A physics model does not care about this; it is solving equations that do not know what year it is. A learned model is an extremely sophisticated argument that the future will rhyme with 1979 through 2018. Every year that argument gets slightly weaker, and nothing in the architecture will tell you when it breaks.

Then there is the tail. The events that kill people are not the ones in the middle of the distribution — they are the once-in-forty-years flood, the storm that stalls, the heat dome nobody has a template for. GenCast's own authors are candid that resolving intrinsically rare, high-impact events requires enormous ensembles, with compute cost climbing steeply as the event gets rarer. The speed advantage gets eaten precisely where you most wanted to spend it. Cheap forecasts for ordinary Tuesdays. Expensive, uncertain forecasts for the day that matters.

Both of those are real technical caveats, and both will probably get better. Which brings me to the two that will not.

The first I have to argue for carefully, because the easy version of it is wrong. The easy version says better forecasts don't save lives. They plainly do. In 1970 the Bhola cyclone killed on the order of 300,000 people in what was then East Pakistan, a coast with 44 shelters and essentially no preparedness plan. Cyclone Sidr hit Bangladesh in 2007 and killed roughly 4,000. That is one of the great humanitarian achievements of the last century, and anyone claiming warning doesn't matter has to explain it.

So look at what Bangladesh actually built. The shelters went from a few hundred to something like 14,000. The Cyclone Preparedness Programme now runs on more than 76,000 volunteers, half of them women, who go door to door with megaphones and physically walk people into the buildings. The forecast was necessary. The forecast was also the cheapest component by an enormous margin. What moved the number from 300,000 to 4,000 was concrete and people, and Bangladesh is not held up as a model because it had the best model.

Now the other end of the income scale. In the United States, where forecast quality is about as good as it gets and every phone is a warning siren, flood deaths run around ninety a year and have sat roughly there across NOAA's thirty-year average. About three-quarters are flash floods. Roughly half are someone driving into water they thought they could cross. That is not a prediction failure. No model improvement addresses it, because the information was already in the car.

It is the shape you find in the after-action report on almost any weather disaster of the last twenty years. Somebody knew. The warning was issued. The model called it. The bulletin went out at 1:14 a.m. to a county where the sirens were down, or where the people at risk were asleep, or where there was nowhere to evacuate to, or where an official read it and decided to wait for the next update.

A model that runs in a minute on one chip is exactly what a national meteorological service with no supercomputer has never been able to afford. Much of the world forecasts on downstream products from Europe and the United States, at someone else's resolution on someone else's schedule. Cheap inference genuinely changes that, and it should. It still isn't the binding constraint, and Bangladesh is why: a sovereign forecast is worth having, and it will be sitting upstream of 14,000 shelters that took forty years and a great deal of money to build.

Knowing has never been the scarce resource. It is the cheap resource, which is exactly why we keep making it cheaper. Prediction scales — train once, serve forever, one chip, one minute, marginal cost approaching zero. The other half does not. You cannot fine-tune a levee. There is no checkpoint you can download that relocates a neighborhood, staffs a volunteer network, or gives a family somewhere to go at two in the morning. That half costs money, takes years, creates enemies, and produces no benchmark you can beat.

Nobody moved money out of levees and into GPUs. Those are different budgets in different institutions, and I can't show you an agency that cut warning dissemination because a model got faster. The asymmetry isn't in appropriations, it's in prestige. Benchmarkable work attracts talent, press, and careers; unbenchmarkable work attracts a line item. That alone explains where the best people in the field spend the next decade, and it doesn't require anybody to have made a decision.

Which brings me to the second thing that will not improve, and it's the one I don't see anyone writing about.

GraphCast and GenCast learned to forecast from ERA5 — decades of global atmospheric state, assembled at public expense by an intergovernmental agency, using the physics stack the learned models are now reported to have beaten. ERA5 is not a corpus somebody scraped. It is the output of a 4D-Var assimilation system fusing observations with short-range forecasts, which is to say it is manufactured by the thing being declared obsolete.

The dependency doesn't end at training. A forecast has to start somewhere, and the somewhere is an analysis: the current state of the atmosphere, estimated by assimilating satellites, buoys, aircraft reports, and the radiosondes that meteorological services still launch twice a day. That is operational numerical weather prediction infrastructure. Nothing in GraphCast produces it. And when these models eventually need retraining on a climate that has moved past 2018 — the fix for the first caveat in this piece — the training data will have to come from a continuously updated reanalysis that only those same institutions know how to make.

So the failure mode isn't exotic. It's a budget meeting where somebody points at a benchmark and observes that the AI is faster and cheaper than the supercomputer, and the supercomputer's line gets trimmed, and some years later there is no ground truth left to learn from. The learned model cannot generate its own future. It is permanently downstream of a commons while being scored as though it replaced one.

The forecast will keep getting better. It will arrive earlier, cleaner, cheaper, with tighter uncertainty bounds, on a model that costs less to run than the meeting where somebody decides whether to act on it. It will land in the same inbox it has always landed in.

The bottleneck was never the storm we couldn't see. It was everything we have to pay for on either side of seeing it — the balloon going up at midnight, and the shelter that has to be standing when the bulletin arrives.

Seeded from

Science News August 2025 — AI weather forecasting outperforms traditional simulations

Science News, August 2025 issue

threaded with