The update
This research update adds completed results for two current models: GPT-6 Astra and Gemini 3.8 Flash.
They were evaluated against the same frozen v4 corpus: 240 short evidence-law items and 60 record-grounded case-file assignments. The overall score gives each half of the benchmark equal weight.
Results
| Model | Overall score | Doctrine | Matter | Fully completed case files |
|---|---|---|---|---|
| Gemini 3.8 Flash | 65.5 | 72.7% | 58.3% | 0.0% |
| GPT-6 Astra | 64.9 | 67.7% | 62.1% | 0.0% |
The four-model result table will appear here once the research release is loaded.
What these scores do—and do not—show
The Doctrine score measures closed-book answers to short evidence-law problems. The Matter score measures work on a supplied record: issue spotting, factual grounding, authority support, and usable deliverables. A fully completed matter assignment must meet every critical requirement; it is deliberately a much stricter measure than partial credit.
This is an attorney-reviewed research release by the project owner, not an independently replicated or official leaderboard release. The results should be read as a reproducible snapshot of this protocol, not a general claim that any model is safe or unsafe for legal practice. Models can produce useful partial work while still missing a legally material issue or authority.
Reproducibility and review
The public evaluator, frozen methodology, and release artifacts are available for scrutiny. The hidden answer key remains controlled so that it can be independently reviewed without contaminating future evaluations.
Explore the interactive EvidenceBench results, inspect the v4 methodology, or volunteer to review a controlled benchmark packet.