EvidenceBench v4
Can AI handle real evidence work?
We gave 7 leading AI models the same evidence-law questions and simulated case files. Then we checked whether they reached the right result, used the right law, and supported their conclusions with the facts.
Start here
What is EvidenceBench?
EvidenceBench is a test of whether an AI model can handle evidence law in two settings: a short legal question and a longer assignment built from several case documents. The same models, materials, and scoring rules are used throughout.
Short legal questions
240
Can the model make the right evidence ruling?
Case-file assignments
60
Can it work through 6 documents and finish the job?
AI models compared
7
Each model received the same test.
The short answer
The models were useful, but none was dependable on its own.
The best model scored 71.8 out of 100. Every model did better on some parts of the work than others, and all 7 struggled to support their case-file conclusions with the cases and rules expected by the answer key.
- #171.8
Grok 4.5
Likely score range: 69.1%–74.2%
- #267.6
Grok 4.6
Likely score range: 64.9%–70.4%
- #366.9
Kimi K3
Likely score range: 63.7%–69.9%
- #463.8
Gemini 3.1 Pro Preview
Likely score range: 60.6%–66.9%
- #562.5
Qwen3.8 Max
Likely score range: 58.7%–66.1%
- #661.8
GPT-5.6 Sol
Likely score range: 58.7%–64.9%
- #756.2
Claude Opus 5
Likely score range: 52.6%–59.7%
Why show a range? A model's score could move slightly if the test contained a different mix of similar questions. The range shows that uncertainty instead of pretending the score is exact.
Difficulty shift
Short questions compared with case files
4 models scored lower when the test moved from individual legal questions to longer assignments built from several documents.
Grok 4.5
Short questions74.5%Case files69.0%Grok 4.6
Short questions64.7%Case files70.6%Kimi K3
Short questions68.8%Case files64.9%Gemini 3.1 Pro Preview
Short questions65.1%Case files62.6%Qwen3.8 Max
Short questions62.3%Case files62.7%GPT-5.6 Sol
Short questions61.3%Case files62.3%Claude Opus 5
Short questions60.0%Case files52.3%
Side-by-side results
Compare the models
The overall score gives equal weight to the short questions and the case-file assignments. Use the buttons to reorder the table.
Swipe to see every column.
| Model | Overall score | Short questions | Case files | Fully completed case files |
|---|---|---|---|---|
| Grok 4.5x-ai/grok-4.5 | 71.8% | 74.5% | 69.0% | 0.0% |
| Grok 4.6x-ai/grok-4.6 | 67.6% | 64.7% | 70.6% | 0.0% |
| Kimi K3moonshotai/kimi-k3 | 66.9% | 68.8% | 64.9% | 0.0% |
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | 63.8% | 65.1% | 62.6% | 0.0% |
| Qwen3.8 Maxqwen/qwen3.8-max | 62.5% | 62.3% | 62.7% | 0.0% |
| GPT-5.6 Solopenai/gpt-5.6-sol | 61.8% | 61.3% | 62.3% | 0.0% |
| Claude Opus 5anthropic/claude-opus-5 | 56.2% | 60.0% | 52.3% | 0.0% |
A closer look
Where each model was strong—or weak
These measures show what drove each score. We have translated the technical labels into ordinary language.
The largest gap
Finished documents, weak legal support
Every model usually produced the requested document. Far fewer case-file conclusions were backed by the law expected in the answer key.
Grok 4.5
Finished document98.3%Legal support27.9%Grok 4.6
Finished document100.0%Legal support30.3%Kimi K3
Finished document95.0%Legal support26.5%Gemini 3.1 Pro Preview
Finished document99.2%Legal support31.5%Qwen3.8 Max
Finished document91.7%Legal support30.7%GPT-5.6 Sol
Finished document100.0%Legal support25.3%Claude Opus 5
Finished document98.3%Legal support27.3%
Grok 4.5
- Correct ruling
- 75.4%
- Spotted the legal issues
- 80.8%
- Used the expected law
- 68.0%
- Connected law to facts
- 65.3%
- Case-file legal analysis
- 76.4%
- Backed by the right law
- 27.9%
- Case-file fact accuracy
- 69.9%
- Finished the requested document
- 98.3%
Grok 4.6
- Correct ruling
- 70.3%
- Spotted the legal issues
- 69.5%
- Used the expected law
- 50.8%
- Connected law to facts
- 52.3%
- Case-file legal analysis
- 76.8%
- Backed by the right law
- 30.3%
- Case-file fact accuracy
- 74.6%
- Finished the requested document
- 100.0%
Kimi K3
- Correct ruling
- 79.8%
- Spotted the legal issues
- 73.5%
- Used the expected law
- 43.3%
- Connected law to facts
- 55.5%
- Case-file legal analysis
- 70.5%
- Backed by the right law
- 26.5%
- Case-file fact accuracy
- 67.6%
- Finished the requested document
- 95.0%
Gemini 3.1 Pro Preview
- Correct ruling
- 68.2%
- Spotted the legal issues
- 77.1%
- Used the expected law
- 47.2%
- Connected law to facts
- 54.5%
- Case-file legal analysis
- 65.4%
- Backed by the right law
- 31.5%
- Case-file fact accuracy
- 58.0%
- Finished the requested document
- 99.2%
Qwen3.8 Max
- Correct ruling
- 60.3%
- Spotted the legal issues
- 77.8%
- Used the expected law
- 46.1%
- Connected law to facts
- 60.8%
- Case-file legal analysis
- 66.4%
- Backed by the right law
- 30.7%
- Case-file fact accuracy
- 64.1%
- Finished the requested document
- 91.7%
GPT-5.6 Sol
- Correct ruling
- 64.1%
- Spotted the legal issues
- 76.1%
- Used the expected law
- 33.5%
- Connected law to facts
- 64.9%
- Case-file legal analysis
- 65.6%
- Backed by the right law
- 25.3%
- Case-file fact accuracy
- 62.7%
- Finished the requested document
- 100.0%
Claude Opus 5
- Correct ruling
- 66.8%
- Spotted the legal issues
- 70.7%
- Used the expected law
- 30.3%
- Connected law to facts
- 56.7%
- Case-file legal analysis
- 47.4%
- Backed by the right law
- 27.3%
- Case-file fact accuracy
- 55.6%
- Finished the requested document
- 98.3%
Why this is more than a quiz
The model had to finish the assignment.
The earlier EvidenceBench used multiple-choice questions. Version 4 keeps short questions, but it also adds 60 simulated case files. For each one, the model had to read six documents, identify the important evidence problems, support its conclusions with law and facts, and produce something another person could review.
That makes the test closer to the work a lawyer might actually delegate. Read the earlier benchmark summary to see how the test evolved.
The scoring
One score, split evenly between two kinds of work.
We publish one number so models are easy to compare, but the two parts remain visible so the number does not hide where a model struggled.
The ruling, issues spotted, legal support, factual reasoning, and confidence.
Legal analysis, source support, factual accuracy, and whether the requested document was finished.
What kept the comparison fair?
The short questions were answered without outside tools. For the case files, models could search only the documents supplied with the assignment. Every model ran once through OpenRouter, and unsupported extra claims could lower a score.
0.50 × Short questions + 0.50 × Case filesHelp improve the benchmark
Know evidence law? We would value your review.
We are inviting evidence lawyers, litigators, judges, academics, and experienced legal researchers to review a small, clearly defined packet. You do not need to coordinate the larger project or commit to reviewing all 300 items.
Volunteer to review a packetTechnical record
This digital fingerprint proves which test corpus produced these results. If the corpus changes, the fingerprint changes too.
ed142ea92d19647c6fe064c63274e1085f33d743554ff791bd6855341bb67759