Skip to content
Legal AI Laboratory

Benchmarks

Evaluating AI systems for legal work.

“Building a Legal AI system is only half the problem. We also need to know whether it works.”

Building better Legal AI requires more than convincing demonstrations. It requires systematic evaluation — and the Laboratory treats that evaluation as research infrastructure, not a leaderboard.

01 · Why Evaluate?

Plausible is not the same as useful.

A legal AI output may sound plausible, be well written, and contain technically correct statements — and still fail to perform the actual legal task. Evaluation has to look past surface-level quality.

  • Did the system identify the relevant issue?
  • Did it use the provided information correctly?
  • Did it miss important facts?
  • Did it reach a supportable conclusion?
  • Did it follow the requested workflow?
  • Did it identify uncertainty?
  • Could an attorney inspect and rely on the work product?

Potential questions — not a finalized or universal standard

Domains

  • Patent Practice
  • Legal Research
  • Litigation
  • Legal Operations
  • Compliance

Technical Areas

  • Software
  • AI / Machine Learning
  • Networking
  • Semiconductor
  • Mechanical
  • Electrical

Status

  • Designing
  • Active
  • Evaluating
  • Published
  • Archived

02 · What We Measure

What We Measure

Dimensions the Laboratory may use when evaluating systems — placeholders for an eventual formal rubric, not finalized methodology.

  1. 01

    Accuracy

    Does the output correctly address the task?

  2. 02

    Completeness

    Does it identify the material issues, facts, or considerations?

  3. 03

    Reasoning / Analysis

    Does the analysis appropriately connect facts, rules, and conclusions?

  4. 04

    Consistency

    Does the system produce reasonably stable results across comparable inputs?

  5. 05

    Evidence / Support

    Are conclusions supported by the relevant source material?

  6. 06

    Workflow Compliance

    Did the system follow the defined workflow?

  7. 07

    Human Reviewability

    Can an attorney efficiently inspect and evaluate the output?

  8. 08

    Failure Severity

    When the system fails, how consequential is the failure?

03 · Benchmark Suite

Benchmark Suite

A structured, evolving collection of benchmark projects — not a fixed set.

Additional benchmark projects will appear here as the suite expands.

04 · Evaluation Method

Evaluation Method

A conceptual view of the evaluation loop. The exact methodology will be based on the Laboratory's actual benchmark implementation as it is finalized.

General

  1. Benchmark Task
  2. AI System
  3. Generated Work Product
  4. Comparison / Evaluation
  5. Scoring
  6. Failure Analysis

Where Human Evaluation Is Involved

  1. Generated Work Product
  2. Human Evaluation
  3. Structured Rubric
  4. Score + Observations

05 · Results

Results — Coming Soon

No benchmark runs have been completed yet. The presentation layer below is ready for real results — every value shown here is an illustrative placeholder, not a finding.

Score Comparison (Illustrative)

  • System A
  • System B
  • System C

Category Performance (Illustrative)

Illustrative category performance table, values pending
RowSoftwareAI / Machine LearningNetworking
System A
System B

06 · Failure Analysis

A score is a summary, not an explanation.

A benchmark shouldn't only tell us that a system scored well or poorly — it should help explain why. Failure analysis is a first-class part of the Laboratory's benchmark design, not an afterthought.

  • Missed issue
  • Incorrect reasoning
  • Unsupported conclusion
  • Hallucination
  • Incomplete analysis
  • Context failure
  • Instruction-following failure
  • Workflow failure
  • Inconsistent output
  • Human-review concern

Provisional categories — not yet formally adopted

Where a benchmark result surfaces an unexpected pattern, the Laboratory documents it as a Field Note — connecting the result back to the research loop that produced it.

  1. Benchmark Result
  2. Unexpected Pattern
  3. Field Note
  4. New Experiment

07 · Limitations & What Comes Next

This is infrastructure, not a verdict.

No single score can fully measure legal AI capability. As the benchmark suite and methodology develop, this section will document known limitations, methodology changes, and future benchmark sets.

The Laboratory is not trying to prove that Legal AI works. It is trying to find out where it works, where it fails, how badly it fails, and what to investigate next.