Benchmarks
Evaluating AI systems for legal work.
“Building a Legal AI system is only half the problem. We also need to know whether it works.”
Building better Legal AI requires more than convincing demonstrations. It requires systematic evaluation — and the Laboratory treats that evaluation as research infrastructure, not a leaderboard.
01 · Why Evaluate?
Plausible is not the same as useful.
A legal AI output may sound plausible, be well written, and contain technically correct statements — and still fail to perform the actual legal task. Evaluation has to look past surface-level quality.
- Did the system identify the relevant issue?
- Did it use the provided information correctly?
- Did it miss important facts?
- Did it reach a supportable conclusion?
- Did it follow the requested workflow?
- Did it identify uncertainty?
- Could an attorney inspect and rely on the work product?
Potential questions — not a finalized or universal standard
Domains
- Patent Practice
- Legal Research
- Litigation
- Legal Operations
- Compliance
Technical Areas
- Software
- AI / Machine Learning
- Networking
- Semiconductor
- Mechanical
- Electrical
Status
- Designing
- Active
- Evaluating
- Published
- Archived
02 · What We Measure
What We Measure
Dimensions the Laboratory may use when evaluating systems — placeholders for an eventual formal rubric, not finalized methodology.
- 01
Accuracy
Does the output correctly address the task?
- 02
Completeness
Does it identify the material issues, facts, or considerations?
- 03
Reasoning / Analysis
Does the analysis appropriately connect facts, rules, and conclusions?
- 04
Consistency
Does the system produce reasonably stable results across comparable inputs?
- 05
Evidence / Support
Are conclusions supported by the relevant source material?
- 06
Workflow Compliance
Did the system follow the defined workflow?
- 07
Human Reviewability
Can an attorney efficiently inspect and evaluate the output?
- 08
Failure Severity
When the system fails, how consequential is the failure?
03 · Benchmark Suite
Benchmark Suite
A structured, evolving collection of benchmark projects — not a fixed set.
Additional benchmark projects will appear here as the suite expands.
04 · Evaluation Method
Evaluation Method
A conceptual view of the evaluation loop. The exact methodology will be based on the Laboratory's actual benchmark implementation as it is finalized.
General
- Benchmark Task
- AI System
- Generated Work Product
- Comparison / Evaluation
- Scoring
- Failure Analysis
Where Human Evaluation Is Involved
- Generated Work Product
- Human Evaluation
- Structured Rubric
- Score + Observations
05 · Results
Results — Coming Soon
No benchmark runs have been completed yet. The presentation layer below is ready for real results — every value shown here is an illustrative placeholder, not a finding.
Score Comparison (Illustrative)
- System A—
- System B—
- System C—
Category Performance (Illustrative)
| Row | Software | AI / Machine Learning | Networking |
|---|---|---|---|
| System A | — | — | — |
| System B | — | — | — |
06 · Failure Analysis
A score is a summary, not an explanation.
A benchmark shouldn't only tell us that a system scored well or poorly — it should help explain why. Failure analysis is a first-class part of the Laboratory's benchmark design, not an afterthought.
- Missed issue—
- Incorrect reasoning—
- Unsupported conclusion—
- Hallucination—
- Incomplete analysis—
- Context failure—
- Instruction-following failure—
- Workflow failure—
- Inconsistent output—
- Human-review concern—
Provisional categories — not yet formally adopted
Where a benchmark result surfaces an unexpected pattern, the Laboratory documents it as a Field Note — connecting the result back to the research loop that produced it.
- Benchmark Result
- Unexpected Pattern
- Field Note
- New Experiment
07 · Limitations & What Comes Next
This is infrastructure, not a verdict.
No single score can fully measure legal AI capability. As the benchmark suite and methodology develop, this section will document known limitations, methodology changes, and future benchmark sets.
The Laboratory is not trying to prove that Legal AI works. It is trying to find out where it works, where it fails, how badly it fails, and what to investigate next.