Five frontier LLMs disagree on 67% of 1k real-world fact-check claims
Lenz Research ยท Snapshot v1.0 ยท data as of May 21, 2026
of real fact-checks, top AI models don't agree on the answer.
Jordanov, Kosta ยท Lenz Research ยท kosta@lenz.io
We presented 1,000 recent real user claims to the five top frontier LLMs and asked each one for a verdict. These aren't benchmark items with public answer keys โ they're claims real users submitted for verification to a fact-checking platform. Only one verdict bucket can be correct per claim, so any disagreement among the panel means at least one model's verdict is label-inconsistent under this 4-bucket rubric (True / Mostly True / Misleading / False). On 67% of claims, the panel splits.
On 67% of claims (672 / 1,000; 95% CI: 64โ70%), the frontier panel doesn't agree โ at least one model dissents from the majority verdict, or no strict majority forms at all. The breakdown:
For each claim we looked at the five frontier verdicts and asked: did at least three pick the same answer (a strict majority)? If yes, how many of the remaining models dissented? If no clear majority emerged at all โ verdicts split across three or four different buckets โ the claim falls in the Models split, no majority row. Most of these claims are unlikely to appear in any training corpus with a gold label attached โ there's no canonical answer key to pattern-match against, no benchmark leaderboard to anchor to.