Skip to content
Legal AI Laboratory
← We Test

Test 001 · Task Capability

Invention Analysis Benchmark

Testing whether an AI system can perform invention analysis across diverse technical disclosures.

Patent Prosecution · Invention Analysis · AI System · Multi-Agent AI

Aug 2026

01 · Question

Can an AI system perform invention analysis across different invention disclosures?

Invention analysis is an early stage of patent prosecution in which a technical disclosure is analyzed to identify the underlying technical problems, inventive concepts, technical features, and potential patentability positions. The quality of invention analysis shapes downstream claim strategy and patent drafting.

This test evaluates whether a structured AI workflow can perform invention analysis consistently across different technologies.

02 · Method

Test Set

The test set includes 12 publicly available technical disclosures selected to represent a range of invention types and technology domains. It is intended to evaluate the reasoning capability of the Invention Analysis workflow rather than its ability to summarize technical publications.

Table 1 · Test Set by Technology Domain

Technology DomainCases
Software & Cloud Systems2
Networking & Distributed Systems3
Semiconductors2
AI / Graphics / Machine Learning2
Mechanical & Electrical Systems3
Total12

Each benchmark disclosure describes a concrete technical invention, contains sufficient implementation detail, is not written in patent claim language, and represents the type of technical material that may serve as the starting point for patent prosecution.

System

The system under test is the structured, multi-agent Invention Analysis workflow described in We Build → Workflows.

The Invention Analysis workflow proceeds through four stages:

Stage 1 — Invention Summary: identifies technical problems, technical approach, and technical advantages.

Stage 2 — Inventive Concepts: identifies and abstracts underlying inventive concepts.

Stage 3 — Technical Feature Analysis & Classification: identifies technical features and classifies them as distinguishing, supporting, or dependent-claim features.

Stage 4 — Patentability Evaluation: evaluates technical significance, disclosure support, prior-art risk, overall patentability, and relative importance of distinguishing features.

Evaluation Method

Each benchmark disclosure is evaluated against a frozen Gold Standard prepared through attorney review and iterative refinement.

  1. Independent understanding. A patent attorney first reviews the disclosure to establish an independent understanding of the invention and its technical contribution.
  2. Initial AI analysis. The Invention Analysis workflow generates an initial Invention Analysis Report.
  3. Attorney review and revision. The generated report is reviewed and refined through the revision workflow until it reaches the quality expected of an experienced patent attorney.
  4. Gold Standard. The final reviewed report is frozen as the Gold Standard.
  5. First-pass evaluation. The Invention Analysis workflow then takes the same benchmark disclosure to generate a first-pass Invention Analysis Report. This output is compared against the frozen Gold Standard using the benchmark scoring rubric below.

Benchmark Scoring Rubric

The evaluation rubric scores the system across four analytical stages and one cross-stage consistency measure.

Table 2 · Benchmark Scoring Rubric

ItemsWeight
Stage 1: Invention Summary20
Stage 2: Inventive Concepts20
Stage 3: Technical Feature Analysis & Classification35
Stage 4: Patentability Evaluation20
Cross-stage Consistency5
Total100

The largest weighting is assigned to technical feature analysis and classification because identifying and characterizing the technical features that support downstream claim strategy is a central objective of the workflow.

03 · Results & Discussion

Results

Across the 12 disclosures, the first-pass reports achieved scores ranging from 84 to 96, with an average score of 88.5/100.

Table 3 · Scoring Results

IDDomainDifficultyStage 1Stage 2Stage 3Stage 4Cross-stageTotal
IA-001Software & Cloud SystemsLow19173119591
IA-002Software & Cloud SystemsHigh20163017588
IA-003Networking & Distributed SystemsMedium20172918589
IA-004Networking & Distributed SystemsMedium-High18172918486
IA-005Networking & Distributed SystemsHigh19163118589
IA-006SemiconductorsMedium19183218592
IA-007SemiconductorsHigh20183419596
IA-008AI / Graphics / Machine LearningLow-Medium20152917586
IA-009AI / Graphics / Machine LearningHigh20142817584
IA-010Mechanical & Electrical SystemsMedium19153019487
IA-011Mechanical & Electrical SystemsHigh19143118486
IA-012Mechanical & Electrical SystemsHigh19163117588

The results indicate that a structured AI workflow can produce reasonably consistent invention analyses across substantially different technical domains, while also revealing meaningful variation between cases.

Discussion

Four observations follow.

  1. Cross-domain performance. The benchmark spans software, networking, semiconductor, AI/graphics, mechanical, and electrical systems, so the workflow was evaluated against substantially different technical vocabularies and invention structures rather than a single narrow domain.
  2. Feature analysis. Technical Feature Analysis & Classification carries the largest portion of the scoring rubric at 35 points, reflecting the importance of distinguishing core technical features from supporting and dependent-claim features.
  3. Case difficulty. Scores vary across cases classified as low, medium, and high difficulty, suggesting that aggregate benchmark performance alone is insufficient to characterize system behavior — individual failure modes and case characteristics also matter.
  4. First-pass evaluation. The benchmark intentionally evaluates first-pass output against a frozen Gold Standard. Revision cycles are used to construct the Gold Standard but are not used to improve the output being scored.

04 · Limitations

This benchmark is an initial evaluation rather than a comprehensive measure of legal reasoning capability.

First, the test set contains only 12 disclosures. Those disclosures are publicly available technical publications rather than confidential invention disclosures. Published technical papers may emphasize implementation details differently from real inventor disclosures.

Second, the benchmark does not evaluate claim drafting, specification drafting, prior-art searching, or ultimate legal conclusions.

Third, the benchmark does not include life sciences or chemical technologies.

05 · Conclusion

This test provides an initial measurement of whether a structured AI workflow can perform invention analysis across diverse technical disclosures.

Across 12 cases, the system achieved an average benchmark score of 88.5/100, with scores ranging from 84 to 96. The results suggest that structured, multi-agent workflows can produce useful first-pass invention analyses across different technology domains.

Additionally, the benchmark illustrates an important distinction: evaluating legal AI is not simply a question of whether a model produces a plausible answer. The evaluation requires a defined task, a controlled test set, an explicit methodology, a reference standard, and a reproducible scoring framework.

Finally, future iterations could expand the test set to additional technical domains, including medical devices, robotics, security, video coding, autonomous systems, and other areas encountered in patent practice.

Related

Related work will be linked here as it is published.