Back to Objection AcademyPublic legal AI research

EvidenceBench v4

Can AI handle real evidence work?

We gave 7 leading AI models the same evidence-law questions and simulated case files. Then we checked whether they reached the right result, used the right law, and supported their conclusions with the facts.

Version 4.0.0-research

Start here

What is EvidenceBench?

EvidenceBench is a test of whether an AI model can handle evidence law in two settings: a short legal question and a longer assignment built from several case documents. The same models, materials, and scoring rules are used throughout.

Short legal questions

240

Can the model make the right evidence ruling?

Case-file assignments

60

Can it work through 6 documents and finish the job?

AI models compared

7

Each model received the same test.

The short answer

The models were useful, but none was dependable on its own.

The best model scored 71.8 out of 100. Every model did better on some parts of the work than others, and all 7 struggled to support their case-file conclusions with the cases and rules expected by the answer key.

  1. #1

    Grok 4.5

    71.8

    Likely score range: 69.1%–74.2%

  2. #2

    Grok 4.6

    67.6

    Likely score range: 64.9%–70.4%

  3. #3

    Kimi K3

    66.9

    Likely score range: 63.7%–69.9%

  4. #4

    Gemini 3.1 Pro Preview

    63.8

    Likely score range: 60.6%–66.9%

  5. #5

    Qwen3.8 Max

    62.5

    Likely score range: 58.7%–66.1%

  6. #6

    GPT-5.6 Sol

    61.8

    Likely score range: 58.7%–64.9%

  7. #7

    Claude Opus 5

    56.2

    Likely score range: 52.6%–59.7%

Why show a range? A model's score could move slightly if the test contained a different mix of similar questions. The range shows that uncertainty instead of pretending the score is exact.

Difficulty shift

Short questions compared with case files

4 models scored lower when the test moved from individual legal questions to longer assignments built from several documents.

  1. Grok 4.5

    Short questions74.5%
    Case files69.0%
  2. Grok 4.6

    Short questions64.7%
    Case files70.6%
  3. Kimi K3

    Short questions68.8%
    Case files64.9%
  4. Gemini 3.1 Pro Preview

    Short questions65.1%
    Case files62.6%
  5. Qwen3.8 Max

    Short questions62.3%
    Case files62.7%
  6. GPT-5.6 Sol

    Short questions61.3%
    Case files62.3%
  7. Claude Opus 5

    Short questions60.0%
    Case files52.3%

Side-by-side results

Compare the models

The overall score gives equal weight to the short questions and the case-file assignments. Use the buttons to reorder the table.

Swipe to see every column.

ModelOverall scoreShort questionsCase filesFully completed case files
Grok 4.5x-ai/grok-4.571.8%74.5%69.0%0.0%
Grok 4.6x-ai/grok-4.667.6%64.7%70.6%0.0%
Kimi K3moonshotai/kimi-k366.9%68.8%64.9%0.0%
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview63.8%65.1%62.6%0.0%
Qwen3.8 Maxqwen/qwen3.8-max62.5%62.3%62.7%0.0%
GPT-5.6 Solopenai/gpt-5.6-sol61.8%61.3%62.3%0.0%
Claude Opus 5anthropic/claude-opus-556.2%60.0%52.3%0.0%

A closer look

Where each model was strong—or weak

These measures show what drove each score. We have translated the technical labels into ordinary language.

The largest gap

Finished documents, weak legal support

Every model usually produced the requested document. Far fewer case-file conclusions were backed by the law expected in the answer key.

  1. Grok 4.5

    Finished document98.3%
    Legal support27.9%
  2. Grok 4.6

    Finished document100.0%
    Legal support30.3%
  3. Kimi K3

    Finished document95.0%
    Legal support26.5%
  4. Gemini 3.1 Pro Preview

    Finished document99.2%
    Legal support31.5%
  5. Qwen3.8 Max

    Finished document91.7%
    Legal support30.7%
  6. GPT-5.6 Sol

    Finished document100.0%
    Legal support25.3%
  7. Claude Opus 5

    Finished document98.3%
    Legal support27.3%

Grok 4.5

Correct ruling
75.4%
Spotted the legal issues
80.8%
Used the expected law
68.0%
Connected law to facts
65.3%
Case-file legal analysis
76.4%
Backed by the right law
27.9%
Case-file fact accuracy
69.9%
Finished the requested document
98.3%

Grok 4.6

Correct ruling
70.3%
Spotted the legal issues
69.5%
Used the expected law
50.8%
Connected law to facts
52.3%
Case-file legal analysis
76.8%
Backed by the right law
30.3%
Case-file fact accuracy
74.6%
Finished the requested document
100.0%

Kimi K3

Correct ruling
79.8%
Spotted the legal issues
73.5%
Used the expected law
43.3%
Connected law to facts
55.5%
Case-file legal analysis
70.5%
Backed by the right law
26.5%
Case-file fact accuracy
67.6%
Finished the requested document
95.0%

Gemini 3.1 Pro Preview

Correct ruling
68.2%
Spotted the legal issues
77.1%
Used the expected law
47.2%
Connected law to facts
54.5%
Case-file legal analysis
65.4%
Backed by the right law
31.5%
Case-file fact accuracy
58.0%
Finished the requested document
99.2%

Qwen3.8 Max

Correct ruling
60.3%
Spotted the legal issues
77.8%
Used the expected law
46.1%
Connected law to facts
60.8%
Case-file legal analysis
66.4%
Backed by the right law
30.7%
Case-file fact accuracy
64.1%
Finished the requested document
91.7%

GPT-5.6 Sol

Correct ruling
64.1%
Spotted the legal issues
76.1%
Used the expected law
33.5%
Connected law to facts
64.9%
Case-file legal analysis
65.6%
Backed by the right law
25.3%
Case-file fact accuracy
62.7%
Finished the requested document
100.0%

Claude Opus 5

Correct ruling
66.8%
Spotted the legal issues
70.7%
Used the expected law
30.3%
Connected law to facts
56.7%
Case-file legal analysis
47.4%
Backed by the right law
27.3%
Case-file fact accuracy
55.6%
Finished the requested document
98.3%

Why this is more than a quiz

The model had to finish the assignment.

The earlier EvidenceBench used multiple-choice questions. Version 4 keeps short questions, but it also adds 60 simulated case files. For each one, the model had to read six documents, identify the important evidence problems, support its conclusions with law and facts, and produce something another person could review.

That makes the test closer to the work a lawyer might actually delegate. Read the earlier benchmark summary to see how the test evolved.

The scoring

One score, split evenly between two kinds of work.

We publish one number so models are easy to compare, but the two parts remain visible so the number does not hide where a model struggled.

50%Short legal questions

The ruling, issues spotted, legal support, factual reasoning, and confidence.

50%Case-file assignments

Legal analysis, source support, factual accuracy, and whether the requested document was finished.

What kept the comparison fair?

The short questions were answered without outside tools. For the case files, models could search only the documents supplied with the assignment. Every model ran once through OpenRouter, and unsupported extra claims could lower a score.

0.50 × Short questions + 0.50 × Case files

Help improve the benchmark

Know evidence law? We would value your review.

We are inviting evidence lawyers, litigators, judges, academics, and experienced legal researchers to review a small, clearly defined packet. You do not need to coordinate the larger project or commit to reviewing all 300 items.

Volunteer to review a packet
Technical record

This digital fingerprint proves which test corpus produced these results. If the corpus changes, the fingerprint changes too.

ed142ea92d19647c6fe064c63274e1085f33d743554ff791bd6855341bb67759