Entities, events, and nothing else
The model has no frames, no grid of pixels, no fixed time step. It represents San Francisco as a large number of entities — fire stations, engines, ambulances, utility feeders, neighborhoods — each with its own state, and the only thing that happens is one entity sending another an event: a call arriving, a unit being dispatched, a feeder losing or regaining power. Nothing is computed for an entity until an event reaches it. A quiet street between calls costs nothing, which is why the full San Francisco model runs twelve days of city time in about 78 seconds on a single processor core.
What's fixed and what's learned
We draw a hard line between the structure of the world and its dynamics. Structure is engineered: which entities exist, what state each one carries, and how they can reach each other. For San Francisco that structure comes from the city’s own topology — fire stations, battalions, feeders, buildings — and is fixed before training starts. Dynamics are learned: how often a district generates a given kind of call, how a unit type tends to respond, how long a crew takes to restore service. Each of those is its own small model, fitted only to the events its part of the world produced.
This is a deliberate, narrower claim than learning the structure itself. Inducing entity types or connections from raw data is real future work, not something this model does today. What it does today is closer to how a robotics or game environment is built: the state space is ours, the behavior inside it is the city’s own.
Training inside the run
Training happens inside the same event-by-event replay used to generate the city, not as a separate offline step. As the model replays recorded history, every observed event becomes a training example for the small model that produced it, fitted to make what actually happened, at the time it happened, as probable as possible. No one labels anything by hand. Each learned piece is tens of thousands of parameters, trained only on its own slice of the city, and the pieces compose on one shared clock at run time.
Two details matter for honesty here. First, a learned piece doesn’t start from nothing: it is warm-started from the real history leading up to the moment it takes over, because a model that has only ever seen an empty history behaves strangely once it does start predicting — we found this the hard way early on, and every learned source now carries real context forward instead of starting cold. Second, the model conditions on the signals we give it: grid status, weather, time of day, calendar. Known confounders like these are given to the model as ordinary inputs, rather than left for it to guess at.
Where the model sits
An event-driven world model is not a new place to keep the facts. It is one participant in a larger arrangement, and it helps to say exactly which one. We keep four worlds distinct, and each answers a different question.
- The observed world is what sources reported, preserved as reported. Here that means the fire department’s call records and the outage record. It is the one layer whose only job is fidelity.
- The believed world is the current, attributed account of how things stand: estimates, assessments, interpretations. It can be wrong and cheaply corrected, because it is built from the layer beneath it.
- The possible world is a claim about what might happen or might have happened: a forecast, a what-if, a rewind. It is always marked as possible, and never mistaken for what was observed.
- The committed world is what an organization actually decided and set in motion.
The boundary we hold to is one sentence: a model may represent the world; it may never become the record of the world. A model is judged by how useful it is, and improves by being revised. A record is judged by fidelity, and survives by refusing revision. When one system plays both roles, usefulness starts editing fidelity. So the world model here is a source and a consumer of the record: it learns from what was recorded, and what it predicts goes back as a claim, with the model and its version as the author, to be graded as outcomes arrive. Its internal state is compressed, revisable and continuously refitted, which are virtues in a model and disqualifications in a record.
That is why every result on this site is stated as relative to what the model has learned. The rewinds in the demo live in the possible world. They are claims about what the model expects, not entries in the history of San Francisco.
One truth. Many projections. Facts live once; every view of them, the model included, is generated from them. The architecture papers set out the doctrine.
Rewinds: replay into a possible world
A rewind has three steps, and the first matters most. It reconstructs what was knowable at the chosen moment from the record: that step reconstructs and does not hypothesize. Then it changes exactly one thing. Then it replays the rest forward as a possible world. Because events only affect what comes causally after them, the model is rolled back to that moment, handed the change, and run forward again, and nothing the change doesn’t reach is recomputed, so nothing about it can differ.
That is what makes a result like Bayview’s — the one San Francisco district the Mission substation doesn’t serve — preserved exactly, byte for byte, across every rewind we’ve run: not a modeling choice, a structural guarantee of computing causally. Everything outside the change’s reach stays identical; everything inside it is the model’s claim, and is marked as one.
A single hypothesis run is one draw from a model that still has uncertainty in it. We don’t read one run as the answer; the demo and our results run each hypothesis five independent times and report the spread, not just a mean, exactly because a single run can look more or less dramatic than what the model actually believes.
Verified, not asserted
“Bit-identical on one processor or a thousand” is a claim worth distrusting until it’s checked, so we check it on every change to the code. Weight files produced by training the same history sequentially and across multiple processors in parallel are compared byte for byte; the committed record of what happened is compared as a set of rows, not a file order, because two correct schedules can commit the same facts in a different order. Both checks are run as a matter of course, not as a one-time demonstration.
Honest limits
These are not hedges. They are specific, measured gaps between what the model predicts and what the record shows, from our own validation. We would rather a reader find them here than find them for us.
Where the model and the record disagree, we keep the disagreement visible rather than smoothing it over. The elevator and electrical-hazard mismatches below are contradictions preserved: the model’s account sits beside what was recorded, unaveraged and unhidden. Conflict is data; resolution is judgment, and a gap closed quietly would be a gap we could no longer learn from.
On the one real blackout we held out, two classes of call are backwards
We trained on San Francisco’s records through April 2025 and evaluated on 20 December 2025 without ever training on it. For the dominant category, medical calls, the model is close: 183 true against 170–188 generated across five independent runs. But elevator rescues and electrical-hazard calls come out inverted. The real afternoon saw 53 elevator rescues and only 3 electrical-hazard calls; the model generated 2–7 elevator rescues and 28–37 electrical-hazard calls across the same five runs — roughly the opposite shape. The model had learned that outages of this kind, in its training history, tend to produce a storm-driven mix of mostly electrical-hazard calls; 20 December was a calm-day outage at a scale the records never showed on a calm day, and the model generalized to the wrong regime for it.
Two call types are miscalibrated even on ordinary days
On two held-out ordinary days with no blackout, most categories track well — elevator rescues matched exactly, medical calls within 9%. But alarms are systematically undercounted (43–53% of the true count across three runs) and outside-structure fires are systematically overcounted (170–210% of true). These are residual calibration gaps independent of the blackout-specific miss above, and they are the kind of thing more data and a closer look at those two categories should fix, not evidence of a deeper problem with the approach.
The model can flag when it's guessing, but not reliably
We train small ensembles of the demand model and watch how much its members disagree. Disagreement does rise on the held-out blackout day relative to ordinary days, and it does flag two of the categories we now know were wrong — elevator and electrical. But the correlation between disagreement and actual error, measured against the categories we can check, is weak (a Spearman rank correlation of 0.17–0.24 across two variants of the method). It also mis-flags: a category with almost no error ranked as the single most disagreeing, while a category with real, moderate error ranked near the bottom. Treat disagreement as a hint worth a second look, not a reliable detector.
Some operational details are stated assumptions, not learned
A few mechanisms in the model rest on an engineering assumption because no real duration data exists for them — for instance, how long a utility crew spends on a precautionary safety check before a repair job. We pick a defensible, clearly labeled value and move on rather than disguise it as learned. The essay’s own account of what’s next names two more: how utility crews are dispatched and how traffic responds to an outage are both encoded as stated rules today, and are candidates to become learned behavior as richer data arrives.
What we're doing about it
The blackout-day miss is the one we care about most, because it is the whole point of the demonstration: a model is only interesting if it tells you the truth about the day it has never seen. We tried adding a building-level outage-exposure signal aimed directly at the elevator/electrical gap; on the held-out blackout day it did not close the gap, and mildly worsened the electrical over-count. We are reporting that as found, not quietly dropping the attempt. Next, we are looking at richer, class-specific training signal rather than a single added feature, and at whether the model should see more of the covariate’s own history, not just its instantaneous value, when an outage is unlike anything in its training window.
If you want to help close these gaps, or you hold records that would let us check the model against a second city, see working with us.
One more honesty, about lineage. None of the ingredients here is new: append-only records, point-in-time reconstruction, provenance and causal replay are all established ideas. The combination is the contribution — one record that models, replays and answers are all held accountable to.