Fire to Continent

The world runs on events

A world model worth having should learn how the world works by watching it, predict what happens when we act, and be able to tell us why. We think the way to build one is from events, the small discrete things that actually happen, like a breaker tripping or a call coming in, out of which everything larger is made. To see how far the idea goes, we built an event-driven world model of San Francisco from the city’s own records and tested it on a day it had never seen: 20 December 2025, when a fire at a single substation left a third of the city without power.

What no other world model can do

To our knowledge, no other learned world model can do any of these. Ours does all six, on a real city.

  1. Rewind and test your own hypothesis, exactly.

    Roll the world back to any moment, change one thing and run it forward. Everything the change doesn’t reach is preserved byte for byte, so every difference is a consequence of the change.

    San Francisco: the district outside the hypothesis’s reach was identical in every run.

  2. Say why any single event happened.

    Every event carries its chain of causes, and a but-for test shows whether a given cause was necessary.

    San Francisco: a Sunset electrical-hazard call that exists only because of the outage.

  3. Build the macro world from micro events.

    City-wide phenomena emerge from millions of learned micro events, one entity at a time, rather than being modeled from the top down.

    San Francisco: one model spans a substation fire, a third of the city in the dark, and individual emergency calls.

  4. Spend compute only where something happens.

    Cost follows events, not area or frames, so a quiet street costs nothing.

    San Francisco: twelve days of the city in about 78 seconds on one processor.

  5. Give the same answer on one processor or millions.

    Results, including what the model learns, are bit-identical at any scale.

    Verified on every change we make.

  6. Learn from more data on ordinary processors.

    More data adds more small models, each trained where its part of the world runs, rather than a bigger network that only GPUs can train.

    San Francisco: built, trained and run entirely on CPUs.

Predicting events, not pixels

Much of the recent excitement about world models comes from generative models that predict the next frame of video, and the images they produce are genuinely impressive. For the questions we care about, though, pixels are the wrong level of description. The consequences of a decision during a blackout aren’t in what the street looks like. They’re in which calls come in, which units are free and which circuits are live. Those are events. A model that predicts events is predicting at the level where cause and effect actually operate, and it can afford to, because it computes only when something happens. A quiet street costs it nothing.

So we represent the world as a very large number of entities (fire stations, engines, utility crews, the demand in each district), each with its own state. Entities exchange events, and each learns from real records what it tends to do next and when. Nothing larger is programmed in. A blackout, a surge of emergency calls, a fire service stretched thin: these are what you get when enough micro events combine, just as in the real city.

Causal by construction

This construction gives us the property we value most. The model is causal by design. An event can only be produced by events that came before it and can only affect events that come after, and influence passes from entity to entity rather than jumping across the world. When you ask the model why something happened, it doesn’t have to guess. The chain of causes is the model.

It also means you can roll the world back to any moment, pose a hypothesis and run it forward again. Everything the hypothesis doesn’t touch is preserved exactly, so whatever differs is a consequence of your hypothesis and of nothing else. The hypothesis can be as small as a single decision, say one more engine waiting at Station 1, and you can watch its effect spread outward, event by event, until it shows up in the city’s totals.

Complexity comes from the same place. A single city produces millions of events a year. A continent over a season produces billions, and a world that adds people, vehicles and weather runs into the trillions. The engine underneath processes billions of events per second, which puts worlds of that richness within reach while every event in them stays traceable to its causes.

Learning the world piece by piece

Because the world already comes divided into entities, learning can be divided too. Rather than one enormous network trained to absorb everything, each part of the world gets its own model, trained only on the events that part produced. Demand in each district is learned from that district’s calls, and each unit type’s behavior from its own dispatches. The pieces are then composed on a single clock, where they interact the way the real entities do.

Training happens inside the world. As it replays recorded history, each observed event becomes a training example for the component that produced it, and each component is fitted to make what actually happened, when it happened, as probable as possible. No one labels anything. Decomposing the problem this way buys a lot. Components can be small. They can be learned from whatever data exists for them and swapped out when better data arrives. Known mechanisms, such as physics or a published dispatch rule, can sit explicitly alongside learned behavior. And the components learn in parallel wherever they run, with identical results on one processor or millions.

Learning from observation also lets a model recover structure it is never shown. In a simulated utility system, we trained a model on nothing but what an outside observer sees, outage reports and restorations, with no access to the dispatcher’s queue or the crews. It learned the crew operation well enough to reproduce the system’s outcomes in closed loop to within 3 to 13%, and it predicted the effect of mutual-aid crews it had never encountered, tracking the true effect with a correlation of 0.97 to 0.99.

The data corrected us, too. We expected dark traffic signals to slow emergency response during outages. Across more than 370 outages, San Francisco’s records show response times holding steady, if anything slightly faster, so the model learned no such effect. When a model learns from observation, the world gets the final say.

More data, more processors

Every result on this page was produced on ordinary CPUs. We used no GPUs, either to run the world or to train it. A single processor core runs twelve days of San Francisco in about 78 seconds, and retrains the city’s models on fifteen months of history in about twelve minutes.

The difference that matters shows up at scale. A foundation model is one network. The more it learns from, the larger it has to become, and every example passes through all of it, which is why training one takes clusters of GPUs. In an event-driven world, more data brings more of the world with it, not a bigger network. A new region, a longer history or a new kind of entity adds components. Each is small, tens of thousands of parameters, trained on its own events alongside the part of the world it describes. The work grows with the number of events, and it spreads across processors as the world does. Growing from a city to a continent means adding ordinary processors, on hardware that is cheap and everywhere.

GPUs remain a welcome accelerator. They can shorten training for larger components and for searches over many alternatives, and the design takes them on naturally. They make learning faster. Scale comes from the structure of the world itself.

A city as a testbed

Games gave AI research something rare: clean environments with unambiguous outcomes, where ideas could be tested and compared. For world models of real societies, we think cities can play that role. They keep detailed records, they are complex enough to matter, and now and then something extraordinary happens that no model has seen.

At 1:09 pm on Saturday 20 December 2025, a fire broke out at PG&E’s Mission substation in San Francisco, and about 40,000 customers lost power. Firefighters were on scene at 2:22. Within the hour PG&E de-energized further circuits for their safety, and the outage tripled to about 130,000 customers, roughly a third of the city. The fire was out by 6:26 pm, and full restoration took 63 hours.

The city’s records show what that meant on the ground. Between 1 and 4 pm the San Francisco Fire Department answered 48 elevator rescues, where an ordinary Saturday afternoon brings about one. In the 2 pm hour there were 78 emergency incidents against a typical 24, and seventeen of the city’s twenty trucks were out at once.

Figure 1. Customers without power in San Francisco County, every 30 minutes (EAGLE-I, Oak Ridge National Laboratory, CC BY 4.0). The step between 2 and 3 pm is the safety de-energization, minutes after firefighters arrived.
Figure 2. Emergency incidents per hour on 20 December (bars) against the average of the four previous Saturdays (dashed), with elevator rescues in amber. SFFD Calls for Service (DataSF).
3D replay of San Francisco at 3:30 pm on 20 December 2025, with the reported-dark neighborhoods unlit and emergency units traced along streets
Figure 3. The interactive replay at 3:30 pm on 20 December: 168,674 buildings at their measured heights, the neighborhoods reported dark shown unlit, and every engine, truck and ambulance run traced along real streets. Routes are inferred from dispatch and on-scene times.

The step in Figure 1 is the whole problem in miniature. A few firefighters at one address changed the state of a third of the city, and within the hour the city, in its changed state, was sending those same firefighters dozens of calls. Effects ran from the smallest scale to the largest and back again. This is the behavior we want a world model to capture.

What the model learned

We trained the model on San Francisco’s records up to April 2025: 7.4 million unit dispatches going back to 2000, the 15-minute outage record, weather, fire stations and buildings. The city arrives as data; another city would arrive the same way. Across the 413 outages in that history since 2018, the model learned on its own how each kind of emergency responds when a district goes dark. Where most homes are without power, it expects elevator rescues to rise about fifteenfold, traffic collisions about twofold and medical calls by about a third.

December 2025 was held out entirely. At 12:55 pm on the day, a quarter of an hour before the fire, we handed the model the city and let it generate the afternoon itself, from the real grid, weather and calendar.

Rolling back and testing a hypothesis

Then we played the part of someone with a question. We rolled the world back to 3 pm and asked: what if the substation’s service area had been restored at that moment? The model ran the evening forward from there, five independent times, each set against the evening as it actually unfolded.

More than a third of the evening’s incidents change: 36% on average, between 31 and 39% depending on the run. Around 135 incidents per run are prevented outright. In Bayview, the one district the substation doesn’t serve, every incident is preserved exactly in every run. It is the same evening, byte for byte.

The rewind in the interactive replay at 7 pm, with power restored at 3 pm, prevented incidents marked as violet rings
Figure 4. The rewind at 7 pm, with the substation’s service area restored at 3 pm. Customers out falls to about zero (the struck-through figure is what actually happened). Violet rings mark incidents the restoration prevents, and everything else is preserved. The panel summarizes all five runs.

That last result matters most to us. The hypothesis was a large one, but the answer comes back one event at a time: which calls vanish, where, and which engines stay available as a result. Since nothing differs between the two worlds except what the hypothesis caused, each of those differences can be followed back to its source. Here is one, as the model records it.

At 5:59 pm in the Sunset, where 68% of homes were dark, the model’s rate of medical emergencies stood at 1.69 an hour, about half again its normal 1.11. A medical call came in, and an ambulance was sent. Replayed with power restored at 3 pm, the rate at that moment is 1.13 an hour, back to normal, and the call is prevented. The outage caused it, in the precise but-for sense, relative to what the model has learned.

From prediction to planning

A model that predicts the consequences of actions can be used to choose actions. You can roll back as many times as you like and run each alternative against the same history. We tried seven restoration strategies, measuring what each prevents against the customer-hours of early restoration it takes. Restoring everything at 1:30 pm prevents the most, around 167 incidents, and is also the best value per customer-hour, 0.147 per 1,000 customer-hours restored early. Restoring Sunset and Richmond first gets most of the way there, around 140.

Figure 5. Incidents avoided between 1 pm and 6 am relative to what happened, mean of three runs, whiskers showing the range. Labels give incidents avoided per 1,000 customer-hours restored early.

The same search works from the other end, too, starting from a single micro decision. We rolled the world back to 1 pm and put one more engine at Station 1, the first station to answer the substation fire, then let the model run the next three days. In every run the change starts at Station 1. The first time Station 1 needs the extra engine, it answers a call it would otherwise have passed to a neighbor, and within two to five minutes a neighboring station responds differently too, in three of the five runs Station 8, less than a kilometer away. From there the change travels outward station by station, reaching as many as 17 stations up to four kilometers away and altering the response to as many as 48 of the city’s roughly 1,150 incidents. Everything beyond its reach is preserved exactly.

A single engine is a small decision, and its footprint is correspondingly small, but every step of that footprint can be traced back to the one engine that caused it. This is the same machinery that carries a substation fire across a third of the city.

A continent

Raw scale isn’t the hard part. The engine beneath this model has run on millions of processor cores. The hard part was keeping cause and effect intact across scales once a learned world model sits on top, from a regional grid down to a single elevator. San Francisco is our first evidence that it holds.

A storm stretching from Texas to New England is, computationally, mostly quiet ground punctuated by intense activity, and an event-driven model spends its effort where the storm is. Regions can run on separate machines and still give bit-identical answers, so going from a city to a continent is engineering on a proven foundation. And because every what-if is exact, a ripple that travels two thousand miles can be told apart from coincidence. That is what makes planning at continental scale worth doing: moving crews, rescue teams and supplies across state lines, and testing each plan against the world before anyone commits to it.

Our work ahead

Events without precedent are the frontier. The model learned that outages raise elevator rescues, but on 20 December they ran at around 150 times their usual rate, far beyond anything in the historical record, and the model produced only a handful. Electrical hazards show the reverse. In past outages they rose about threefold on calm days and more than twentyfold when a storm was bringing down lines. The records held no severe outage on a calm day, so the model expected the storm pattern; on 20 December, with no storm, the calm-day rate predicts two or three calls, and there were three. Telling apart causes that look alike in the records is part of this work. We are building ensembles of models whose disagreement tells you when a prediction has moved beyond the model’s experience, and learning at several levels at once, from a single entity up to a type, a population and the coarse behavior of a whole region. Richer data will turn the links that rest on stated assumptions today, such as how utility crews are dispatched and how traffic responds, into learned ones. And we are closing the loops that made the day what it was: firefighters triggering de-energization, downed wires drawing crews toward safety work, and restoration in turn shaping where the next call comes from.

Work with us

Work with us

We are looking for partners: utilities and cities with outage, crew and dispatch records, groups with the compute to run continent-wide worlds, and researchers who want world models that can explain themselves.

Notes

Data. SFFD Fire Department Calls for Service (DataSF, 7.4M unit records, 2000–2026). EAGLE-I county outage data (Oak Ridge National Laboratory, CC BY 4.0). NOAA ISD and ERA5 weather. DataSF building footprints, streets, fire stations and neighborhoods. NREL SMART-DS synthetic distribution grid. Event facts from PG&E and news reporting, including the independent Exponent analysis (May 2026).

Method. Learned demand and response models, trained inside the event-driven world on data before May 2025 and evaluated on the held-out December 2025 event. Counterfactual runs preserve everything outside a hypothesis’s reach. Five runs per scenario, three for the strategy search. San Francisco figures are model output unless marked as recorded or reported. The utility-crew result comes from a simulated system. The event engine’s scalability is established; continent-wide world models with regional data are planned with partners.