Incident response and status page design

How AWS and the EIA each handle telling customers whether a fix is retroactive.

On February 28, 2017, Amazon's S3 storage service went down for four hours after an engineer's routine command removed more capacity than intended, taking a piece of the internet with it. AWS's service health dashboard couldn't be updated because the dashboard's status icons were hosted on the S3 infrastructure that had just gone down.

For two hours, the only way Amazon could tell customers what was happening was by editing a banner by hand and posting updates on Twitter. The system meant to explain the incident had failed. A status page is infrastructure that has to work at the moment everything else doesn't.

Public postmortem

What AWS does afterwards is a model for a customer-facing writeup in its public post-event summary. The company has a standing, published commitment to produce a public Post-Event Summary for any incident with sufficiently broad impact, covering the scope of what was affected, the factors that contributed, and the actions taken afterward, and to keep that summary available for a minimum of five years. This stands in contrast to the internal blameless postmortem culture that companies like Google have written about at length: the internal document is written for engineers who need the full causal chain so the same failure doesn't recur, while the public one exists for customers who mostly need three things. What broke, what it touched, and whether it's over. Those are different documents serving different readers, and conflating them produces something too technical for a customer and too sanitized for an engineer.

For a product built on data rather than uptime, though, even AWS's template is missing an important piece. A data provider needs to provide a list of affected numbers and dates before customers can decide whether to trust anything they've already downloaded.

The US Energy Information Administration writes this kind of notice constantly, because weekly and monthly energy data gets revised often. When a stock error affects the Weekly Petroleum Status Report, EIA publishes an errata table naming the region, product, and week for every affected value, alongside the originally published figure, the corrected figure, and the difference between them, so a reader can see how far off the December 8, 2023 propane stock figure for the Midwest region was. Just as important, the notice states plainly whether the fix is retroactive. In that case, EIA said it would not republish the affected weekly editions, but would use the corrected values going forward for any calculation, such as the week-over-week change, four-week averages etc, that depended on them. That answers the question a customer has, which isn't just "what happened" but "is the number I already used still the number I should have used."

EIA's shorter correction notices for its monthly Short-Term Energy Outlook follow a similar structure. A sentence naming the error, a sentence naming the exact date range and magnitude it affected, and a sentence saying when the corrected version will be published. One from December 2025 named a specific misallocation, around 40,000 barrels per day of Permian region production incorrectly counted elsewhere, for a fifteen month window, and it named the date the fix would land.

There's no root cause narrative in it, and it doesn't need one. A customer deciding whether last week's export is still good doesn't need the story of how the mistake happened. One of EIA's revision notices from August 2026, after a text and data mismatch in a weekly report, went a step further and admitted the notice itself should have gone up the same day the error was caught rather than later.


If you do this kind of screening for a living, get in touch. A walkthrough can use a market you actually cover.