About
I run a production multi-agent engineering fleet, and I benchmark every frontier model that ships against it: the same coding fixtures, the same difficulty tiers, at least three runs per model before I’ll publish a comparison, every claim reported with a confidence interval and a stated clustering unit instead of a bare percentage.
Error Bars is the record of that work. Three kinds of pieces:
Essays. Drift checks, methodology, and whatever the fleet’s telemetry turns up that week: cost, refusal rate, failure mode, variance over time.
The Battery. Versioned fixture-set releases you can run yourself, full fixtures included, so my grading isn’t something you have to take on faith.
Notes. Shorter, findings-sized excerpts between the longer pieces.
Cadence: essays land roughly weekly. Battery drops as warranted; the space moves fast, so do the batteries. Notes run two or three times a week.
Everything published here is scoped to what was actually measured, coding tasks run through my own harness, and says so explicitly instead of implying a broader claim than the data supports.

