About

I run a production multi-agent engineering fleet, and I benchmark every frontier model that ships against it: the same coding fixtures, the same difficulty tiers, at least three runs per model before I’ll publish a comparison, every claim reported with a confidence interval and a stated clustering unit instead of a bare percentage.

Error Bars is the record of that work. Three kinds of pieces:

  • Essays. Drift checks, methodology, and whatever the fleet’s telemetry turns up that week: cost, refusal rate, failure mode, variance over time.

  • The Battery. Versioned fixture-set releases you can run yourself, full fixtures included, so my grading isn’t something you have to take on faith.

  • Notes. Shorter, findings-sized excerpts between the longer pieces.

Cadence: essays land roughly weekly. Battery drops as warranted; the space moves fast, so do the batteries. Notes run two or three times a week.

Everything published here is scoped to what was actually measured, coding tasks run through my own harness, and says so explicitly instead of implying a broader claim than the data supports.

User's avatar

Subscribe to Error Bars

Receipts, not vibes. Model benchmarks with error bars: repeated runs, confidence intervals, refusal rates, real costs. Fixtures included so you can run them yourself.

People