Welcome to Error Bars
Receipts, not vibes.
Launch day for a new frontier model is always the same show. Within an hour there are two competing threads running on the same forum. One says the new model is clearly worse than what it replaced, citing a task that went sideways. The other says the first group is overreacting, the model is fine, maybe even better, citing a task that went great. Both sides ran the model once. Neither posted a sample size, a confidence interval, or the fixture they actually used. By the next morning the argument has fully detached from anything either side actually measured, and it’s still going.
I run a benchmark panel against my own production fleet every time a frontier model ships: coding tasks across four difficulty tiers, the same fixtures across every model in the panel, at least three runs before I’ll publish a comparison. Then I do the part the launch-day threads never get around to. I put an error bar on it.
Here’s what that looks like in practice, per the standard laid out in Anthropic’s “Adding Error Bars to Evals” (arXiv:2411.00640). Every fixture set gets resampled, repeated three or more times, grouped by the underlying question so that three attempts at the same fixture don’t get counted as three independent pieces of evidence, and reported with a clustered confidence interval sitting right next to the raw number. Some of the cleanest-looking results in my own library get noticeably less clean once that correction runs. That’s not the method failing. That’s the method doing its job: telling me which of my point estimates I can stand behind and which ones are a coin flip wearing a finding’s clothes.
That’s the claims-contract this publication runs on. Every number I publish here ships with its sample size, its confidence interval, and the unit it was clustered on, which fixture, which task, which repeated question is doing the work underneath the average. Where I’ve built a fixture set specifically for readers to run themselves, in what I’m calling The Battery, the full set ships downloadable alongside the writeup, so my grading isn’t something you have to take on faith either.
Why now. Frontier launches keep arriving faster than anyone is running proper batteries against them, and the discourse around each one keeps filling that gap with vibes. I’d rather publish fewer, slower numbers that come with their own honest uncertainty attached than another same-day take that reads clean and isn’t. The first real piece goes up today alongside this note, or lands within hours of it: multiple timepoints on the model everyone argued about on day one, and what happens to that argument once you put error bars on both sides of it. More follows over the next few days as the batteries keep accumulating data, and at that pace from then on.

