Skip to content
Legal AI Laboratory

We Test

Can AI actually do legal work?

How do we find out? We build platforms, systems, and workflows, test them on defined legal tasks, and examine not only whether they succeed, but where they fail and why.

We Build → We Test → We Document → Tutorials

Legal AI Testing

A cyclical diagram — Question, Test, Observe, Evaluate, Decide, Learn — arranged in a loop, illustrating the Laboratory's testing cycle.

Questioning · Testing · Evaluation · Iteration

01 · The Testing Framework

How do we test legal AI systems?

Legal AI systems can fail at different levels. We test defined legal tasks, the workflows that connect them, the reliability of the resulting system, and the points where attorney judgment remains necessary.

  1. 01

    Task Capability

    Can AI actually perform a defined legal task?

  2. 02

    Workflow Performance

    Does a structured workflow produce a useful result?

  3. 03

    System Reliability

    Does an AI system behave consistently across cases and executions?

  4. 04

    Attorney Judgment

    Where is attorney judgment needed, and how should it be incorporated?

02 · Test Records

What have we tested?

Each test record captures the question, methodology, results, and what we learned.

Test 001· Task Capability

Invention Analysis Benchmark

Testing whether an AI system can perform invention analysis across diverse technical disclosures.

Patent Prosecution · Invention Analysis · AI System · Multi-Agent AI

Test 002· Workflow Performance

Workflow Reliability and LLMs

Testing whether a structured legal workflow remains reliable when its underlying LLM changes.

Workflow Execution · Reliability · LLMs · Multi-Agent AI

Test 003· Attorney Judgment

Attorney Feedback and Revision Loop

Testing whether attorney judgment can be reliably translated into AI revisions.

Attorney Review · Revision · Human-in-the-Loop · Work Products