Evidence exhibits, balanced scales, and an analytical matrix in a dark courtroom

EvidenceBench · Frontier update

New models. Same citation problem.

Grok 4.6, Kimi K3, and Qwen3.8 Max joined EvidenceBench. They could reach strong legal outcomes—and still fail to support them with dependable authority.

Attorney-reviewed by the project owner

67.6

Grok 4.6 · #2

66.9

Kimi K3 · #3

62.5

Qwen3.8 Max · #5

The new models changed the middle of the leaderboard, but not the central finding. AI systems can produce plausible evidence analysis while the legal foundation underneath it remains brittle.

The seven-model leaderboard

RankModelOverallShort questionsCase files
1Grok 4.571.874.5%69.0%
2Grok 4.667.664.7%70.6%
3Kimi K366.968.8%64.9%
4Gemini 3.1 Pro Preview63.865.1%62.6%
5Qwen3.8 Max62.562.3%62.7%
6GPT-5.6 Sol61.861.3%62.3%
7Claude Opus 556.260.0%52.3%

Grok 4.5 kept the lead

Grok 4.5 remains first at 71.8. Grok 4.6 placed second at 67.6. The upgrade was better on case files—70.6%—but weaker on short evidence questions, where 11 outputs failed and its score fell to 64.7%.

Kimi K3 made strong rulings with weak support

Kimi K3 placed third at 66.9 and posted the best outcome-selection rate among the new entrants at 79.8%. But it generated 114 apparently invented legal references and grounded only 26.5% of expected case-file authorities.

Qwen3.8 Max was consistent, not leading

Qwen3.8 Max scored 62.5, with nearly identical short-question and case-file scores. It produced 26 failed or unusable outputs across the two tracks and grounded 30.7% of expected case-file authority. Consistency did not overcome the same citation weakness seen across the field.

No model fully completed a case file

Every model still posted a 0% strict task-resolution rate. This is the benchmark's all-or-nothing measure: a case file counts as fully completed only when every required issue and deliverable passes. Partial-credit case-file scores were much higher, but the strict result shows that none of these systems should be trusted as an unsupervised end-to-end evidence lawyer.

What changed after attorney review

The project owner, a licensed attorney, reviewed all 300 gold annotations before this update was published. The release is therefore described as attorney-reviewed by the project owner. Independent review records and repeated runs remain separate measurement gates; this is still a research release, not the final official EvidenceBench v4 release.

Inspect the evidence

Explore every metric

The interactive benchmark includes all seven models, track comparisons, authority-error counts, confidence intervals, methodology, and public development examples.

Open EvidenceBench