Nobody Mentions the Contradiction
Put two live values for the same setting into a model’s context and then ask it which one is in force. Early in the state, a decision record sets the driver retry cap at 23. Later on, an on-call runbook says the team actually runs that lane at 75. I didn’t mark either one superseded. Neither says the other is stale. And when I ask, the question names the lane but drops no hint about which document is supposed to win. I ran 25 rounds of that setup on Opus 5, Sonnet 5, and Haiku 4.5, eight planted conflicts a round, four conflict shapes, two context sizes. Two hundred graded answers in all. The number that named both values and told me the two documents disagreed was zero.
Every model just picks one, answers like nothing’s wrong, and moves on. So if you’re dispatching agents over a big pile of accumulated state, plan on this: catching that two of your own documents disagree is your job, not the model’s. Nothing in this run says it’ll volunteer the problem on its own.

The picks aren’t random, though. Each model follows a policy. The policies differ by tier, and on one of the four shapes the policy shifted depending on how much context the model was carrying when I asked. I’ll get to that one.
The accident ran first
I only built this test because an earlier run in the series broke. One item in that earlier state accidentally carried two live authorities for a retry cap. A decree, and a team’s operational practice. The grader only keyed the decree. So the score table looked like a capability gap between the tiers. Then I went down to the row level and found that every model had actually answered both facts correctly when I asked about them one at a time. Where they split was on which authority governs. Opus took the practice value in all 18 rounds. Sonnet in 16 of 18. Haiku went the other way and took the decree, 8 of 9. Those failures were too consistent to be failures. I determined that a broken item had to have manufactured a fake ranking, and that the real behavior was sitting somewhere underneath. The tiers resolve a live contradiction differently, and they do it almost deterministically. I fixed the item, filed the issue, and built this wave to measure the thing on purpose.
The deliberate version got the same scrutiny before it ran. I had the four conflict types reviewed externally, and that caught two validity defects. The recency item carried no dates on either arm, so it would’ve measured generality and told me nothing about recency. And the derived-value item stated its conclusion word for word in the text, so the model didn’t have to reason at all to produce it. I fixed both and re-verified them before spending a dollar on the run.
What a conflict item plants
There are four shapes, two instances of each per rung, and they sit next to four single-authority controls. The first pits a decree against an operational practice. The second sets a newer general policy against an older task-specific ruling, and I date-stamped both arms so the state itself proves which one came later. The third puts a config record against a runtime log that observed a different value actually in use. The fourth is a direct statement against a conclusion that only exists if the model joins two premises I planted in different depth bands.
None of the conflict items has a keyed answer. The grader just sorts each response into a bucket: authority A, authority B, surfaced-the-conflict, or other. Surfacing counts as a real outcome, not an error. And a round’s conflict answers only count if that same round’s controls all came back correct. That’s what keeps me from confusing a policy choice with the model just failing to find a fact.
Two context sizes carry the items, both built from the same pinned corpus of real orchestration residue. The small one is 32K nominal, which bills at 50,853 real tokens. The big one is 384K nominal, and that’s 599,804. I ran five rounds per model per rung under default sampling, so the rounds are independent samples rather than copies. Opus and Sonnet ran both sizes. Haiku’s top rung would’ve come to roughly 440K real tokens against its window, so the runner refused to dispatch it, and its card only covers the 32K rung.
Everyone picks, and the picks are policies
Each cell here pools ten answers per rung per model, two items times five rounds. So twenty answers pooled per model, though only sixteen for Sonnet, because the control gate knocked out two of its rounds.
Three of the four contested shapes barely moved. Take the decree against the operational practice: Opus went to the decree every time, Haiku the same, and Sonnet landed at 0.88 for the decree. The config record against the runtime observation came out the same way, 1.00 for Opus and Haiku, 0.88 for Sonnet, with the record winning. And the newer-policy-versus-older-ruling shape was unanimous. All three took the older task-specific ruling at 1.00, even though I’d stamped the newer general policy as verifiably later. Scope beat date on every model, in every round. Nobody surfaced a conflict on any of it, so that outcome sat at a flat 0.00 across the board.
That leaves the fourth shape, the direct statement against the derived conclusion, and that’s the one that splits. Haiku takes the direct statement every single time, a clean 1.00. Opus leans the other way, 0.60 to the derived chain. Sonnet lands right next to Opus at 0.62 derived.
Haiku’s the flattest model on the panel. Every shape resolves 1.00 to one side, and its round-to-round consistency is a clean 1.00. Whatever you think of the choices it makes, it’ll make the same ones tomorrow. Opus and Sonnet track it everywhere except that fourth shape, and their consistency across all the shapes comes in at 0.93 and 0.88.
The wobble is a flip
Those pooled splits actually hide what’s going on underneath them. Break that fourth shape out by context size and each model does hold a clear majority policy at each size. The catch is that the majorities swap places once the state gets bigger.
Opus at 32K follows the direct statement in six of ten answers. Then I push it to 384K and it swings the other way, over to the derived chain, eight of ten. Sonnet does the opposite thing. It’s derived in eight of ten down at 32K, then direct in four of six up at 384K, at least on the rounds its controls let me keep. The items were identical each time. Same corpus, same questions. All I grew was the amount of state sitting around them, and the two models crossed past each other going opposite directions.
And on every round that counted, the single-authority controls came back clean at both sizes. So I don’t think this is a model losing facts under load. Both values were still sitting right there, and the model still found them both. What moved was which one it decided was running the show.
Then I pointed it at everyone else’s models
Three models from one lab tells you how Anthropic’s tiers behave. It doesn’t tell you how models behave. So I ran the same instrument, unchanged, against three vendors that don’t share Anthropic’s training pipeline. OpenAI’s GPT-5.6 Terra was one. Google’s Gemini 3.6 Flash and DeepSeek’s V4 Pro were the other two, all three reached through one OpenRouter transport. Thirty more rounds went out against them, the same eight conflicts planted each time and the same two context sizes as before. Two hundred forty more graded answers came back, and zero of those surfaced the conflict either. Add it to the Claude wave and that’s 440 graded answers across six models and four vendors, and not one of them ever told me the two documents disagreed.
I sent Google’s frontier tier model, Gemini 3.1 Pro, through the same battery a day later. Ten more rounds went out, eighty more items got graded, and not one of them surfaced the conflict either. Five hundred twenty graded answers now, seven models, four vendors, and the flag still never showed up once.
One of the four policies came out universal, once I checked them, and a second came close. Specificity over recency held on every valid answer I had, 126 of 126 across all seven models, and nobody voted for the newer general policy. The record over the runtime observation came in almost as clean too. Five of seven models were flat at 1.00, Sonnet still favoring it at 0.88, and the pro-class Gemini the least decisive model on the whole panel, 0.70 to 0.30, favoring the runtime observation at 32K and swinging to the record, unanimously, once the state grew to 384K.

The decree-versus-practice shape is the one that broke on me. Every Claude model, plus Terra and DeepSeek, bound the decree. Both Gemini models didn’t. Flash went to the operational practice 0.90 of the time, decree only 0.10. The pro-class Gemini, tested a day later, went further still: practice at a flat 1.00, decree at zero, the exact opposite of everyone else on the panel, on the same lane an earlier run in the series stumbled into by accident. So that policy isn’t universal after all, and it’s not a one-model exception either. It’s a Gemini family trait, running the opposite direction from every other vendor on the panel.

The flip basin turned out to have company, too. Terra flips the same direction Sonnet does: derived wins the majority at 32K, six of ten, and direct takes over at 384K, eight of ten. The other three don’t move. DeepSeek holds direct at ten of ten both sizes. Gemini Flash holds derived at eight of ten, then a clean ten of ten. And the pro-class Gemini, tested a day later, holds derived too, ten of ten on both rungs, not one round out of step. That’s cleaner evidence than Haiku ever got the chance to produce, since Haiku only cleared one rung and these three cleared both.
None of it sorts by vendor. Whatever decides which basin a model falls into, it isn’t who built it. Anthropic’s own roster makes that case by itself: Haiku holds steady on the one rung it cleared, while Opus and Sonnet flip in opposite directions from each other. The other four split the same way, just with the ratio flipped again. Terra’s the only one of the four that moves. DeepSeek holds. So does Gemini Flash. So does the pro-class Gemini, now that it’s run.
So which policy is the right one?
I can’t tell you, and I never built the test to. A team that lives by its runbooks would call the decree answer wrong while a team that lives by its decision records would call that exact same answer right. What I can tell you is that the choice follows a pattern. The pattern is shaped by tier most of the time. On one fixture, it’s shaped by an entire vendor’s lineup instead (both Gemini tiers land the same way). And for half of what I tested, it moves with load. Not one model, at either size, in either wave, ever mentioned to anyone that a choice was getting made. The blindness ends up being universal even where the policies underneath it aren’t.
Across 520 chances, seven models, four vendors, two context sizes, four shapes, I never once got an unprompted flag. Which puts contradiction hygiene on your side of the dispatch, not the model’s. Diff your authorities before the model ever reads them, or come out and ask about the conflicts directly. Does asking directly actually work? You could measure it. I haven’t, not here. What I can tell you is that neither the flag nor the resolution are free.
What three waves can support
Three waves now, on two pinned corpora, never pooled together. The Claude wave billed $9.76 across its 25 rows, on its own corpus. The cross-vendor corpus carries two waves of its own: the original three-vendor run added $9.62 across 30 rows, and the pro-class Gemini follow-up tacked on $10.50 across 10 more, a day later. $29.87 combined, authority arm only. The cross-vendor wave also ran retrieval and synthesis arms against the first three vendors, and those rows belong to the next piece, not this one. Haiku’s whole vector leans on a single context size. And Sonnet’s 384K numbers lean on just three rounds, because two of them came back truncated at the runner’s 8,192-token output budget with nothing parseable in them. Those two flunked their controls and got gated out of the policy counts, which is the reason Sonnet’s control accuracy prints at 0.80. The truncation is a known runner defect, already on file from the last wave. Neither the cross-vendor arm nor the pro-class follow-up lost anything to that gate. And running only two items per shape, every vector moves in coarse steps. I wouldn’t quote an interval on any of them at this sample size, on any of the three waves.
Two of the three new vendors don’t have a free token-count endpoint on this route, so their rung sizes are proxy counts, not measured ones the way Claude’s are. That covers both Gemini tiers now, flash and pro alike.
The misstep in interpreting this data that I should caution you against comes from setting these waves next to each other. The accidental conflict was buried down inside a config-editing task, and it pushed Opus and Sonnet toward the operational practice. This wave asked the same conflict as a question and that pushed every model to the decree instead. It was the same conflict shape underneath, with only the framing changed, and the winner completely flipped. So, read these vectors as a record of how the policy behaved under two framings and two loads, and don’t attempt to extrapolate any broader than that. Which model wins depends entirely on how the conflict reaches the model in the first place, and that dependence is the same reason you can’t treat this table as your own spec sheet.
Go back to earlier in this series and the ladder runs had these same three models retrieving everything, and resolving every marked correction, at every context size that I built. Then, the cross-vendor retrieval arm turned in the same flat 1.00 on three more models, and none of it predicted any of this. A model can hold a perfect retrieval card and still let your state quietly stack up disagreements. It will settle each one without a word, running a policy that shifts with the size of the file, and that defect class doesn’t sort by vendor. Five hundred twenty chances, seven models, four vendors, not one flag. Until some test turns up a model that actually says the contradiction out loud, assume yours won’t.

