Test 001 · Task Capability
Invention Analysis Benchmark
Testing whether an AI system can perform invention analysis across diverse technical disclosures.
Patent Prosecution · Invention Analysis · AI System · Multi-Agent AI
Aug 2026
01 · Question
Can an AI system perform invention analysis across different invention disclosures?
Invention analysis is an early stage of patent prosecution in which a technical disclosure is analyzed to identify the underlying technical problems, inventive concepts, technical features, and potential patentability positions. The quality of invention analysis shapes downstream claim strategy and patent drafting.
This test evaluates whether a structured AI workflow can perform invention analysis consistently across different technologies.
02 · Method
Test Set
The test set includes 12 publicly available technical disclosures selected to represent a range of invention types and technology domains. It is intended to evaluate the reasoning capability of the Invention Analysis workflow rather than its ability to summarize technical publications.
Table 1 · Test Set by Technology Domain
| Technology Domain | Cases |
|---|---|
| Software & Cloud Systems | 2 |
| Networking & Distributed Systems | 3 |
| Semiconductors | 2 |
| AI / Graphics / Machine Learning | 2 |
| Mechanical & Electrical Systems | 3 |
| Total | 12 |
Each benchmark disclosure describes a concrete technical invention, contains sufficient implementation detail, is not written in patent claim language, and represents the type of technical material that may serve as the starting point for patent prosecution.
System
The system under test is the structured, multi-agent Invention Analysis workflow described in We Build → Workflows.
The Invention Analysis workflow proceeds through four stages:
Stage 1 — Invention Summary: identifies technical problems, technical approach, and technical advantages.
Stage 2 — Inventive Concepts: identifies and abstracts underlying inventive concepts.
Stage 3 — Technical Feature Analysis & Classification: identifies technical features and classifies them as distinguishing, supporting, or dependent-claim features.
Stage 4 — Patentability Evaluation: evaluates technical significance, disclosure support, prior-art risk, overall patentability, and relative importance of distinguishing features.
Evaluation Method
Each benchmark disclosure is evaluated against a frozen Gold Standard prepared through attorney review and iterative refinement.
- Independent understanding. A patent attorney first reviews the disclosure to establish an independent understanding of the invention and its technical contribution.
- Initial AI analysis. The Invention Analysis workflow generates an initial Invention Analysis Report.
- Attorney review and revision. The generated report is reviewed and refined through the revision workflow until it reaches the quality expected of an experienced patent attorney.
- Gold Standard. The final reviewed report is frozen as the Gold Standard.
- First-pass evaluation. The Invention Analysis workflow then takes the same benchmark disclosure to generate a first-pass Invention Analysis Report. This output is compared against the frozen Gold Standard using the benchmark scoring rubric below.
Benchmark Scoring Rubric
The evaluation rubric scores the system across four analytical stages and one cross-stage consistency measure.
Table 2 · Benchmark Scoring Rubric
| Items | Weight |
|---|---|
| Stage 1: Invention Summary | 20 |
| Stage 2: Inventive Concepts | 20 |
| Stage 3: Technical Feature Analysis & Classification | 35 |
| Stage 4: Patentability Evaluation | 20 |
| Cross-stage Consistency | 5 |
| Total | 100 |
The largest weighting is assigned to technical feature analysis and classification because identifying and characterizing the technical features that support downstream claim strategy is a central objective of the workflow.
03 · Results & Discussion
Results
Across the 12 disclosures, the first-pass reports achieved scores ranging from 84 to 96, with an average score of 88.5/100.
Table 3 · Scoring Results
| ID | Domain | Difficulty | Stage 1 | Stage 2 | Stage 3 | Stage 4 | Cross-stage | Total |
|---|---|---|---|---|---|---|---|---|
| IA-001 | Software & Cloud Systems | Low | 19 | 17 | 31 | 19 | 5 | 91 |
| IA-002 | Software & Cloud Systems | High | 20 | 16 | 30 | 17 | 5 | 88 |
| IA-003 | Networking & Distributed Systems | Medium | 20 | 17 | 29 | 18 | 5 | 89 |
| IA-004 | Networking & Distributed Systems | Medium-High | 18 | 17 | 29 | 18 | 4 | 86 |
| IA-005 | Networking & Distributed Systems | High | 19 | 16 | 31 | 18 | 5 | 89 |
| IA-006 | Semiconductors | Medium | 19 | 18 | 32 | 18 | 5 | 92 |
| IA-007 | Semiconductors | High | 20 | 18 | 34 | 19 | 5 | 96 |
| IA-008 | AI / Graphics / Machine Learning | Low-Medium | 20 | 15 | 29 | 17 | 5 | 86 |
| IA-009 | AI / Graphics / Machine Learning | High | 20 | 14 | 28 | 17 | 5 | 84 |
| IA-010 | Mechanical & Electrical Systems | Medium | 19 | 15 | 30 | 19 | 4 | 87 |
| IA-011 | Mechanical & Electrical Systems | High | 19 | 14 | 31 | 18 | 4 | 86 |
| IA-012 | Mechanical & Electrical Systems | High | 19 | 16 | 31 | 17 | 5 | 88 |
The results indicate that a structured AI workflow can produce reasonably consistent invention analyses across substantially different technical domains, while also revealing meaningful variation between cases.
Discussion
Four observations follow.
- Cross-domain performance. The benchmark spans software, networking, semiconductor, AI/graphics, mechanical, and electrical systems, so the workflow was evaluated against substantially different technical vocabularies and invention structures rather than a single narrow domain.
- Feature analysis. Technical Feature Analysis & Classification carries the largest portion of the scoring rubric at 35 points, reflecting the importance of distinguishing core technical features from supporting and dependent-claim features.
- Case difficulty. Scores vary across cases classified as low, medium, and high difficulty, suggesting that aggregate benchmark performance alone is insufficient to characterize system behavior — individual failure modes and case characteristics also matter.
- First-pass evaluation. The benchmark intentionally evaluates first-pass output against a frozen Gold Standard. Revision cycles are used to construct the Gold Standard but are not used to improve the output being scored.
04 · Limitations
This benchmark is an initial evaluation rather than a comprehensive measure of legal reasoning capability.
First, the test set contains only 12 disclosures. Those disclosures are publicly available technical publications rather than confidential invention disclosures. Published technical papers may emphasize implementation details differently from real inventor disclosures.
Second, the benchmark does not evaluate claim drafting, specification drafting, prior-art searching, or ultimate legal conclusions.
Third, the benchmark does not include life sciences or chemical technologies.
05 · Conclusion
This test provides an initial measurement of whether a structured AI workflow can perform invention analysis across diverse technical disclosures.
Across 12 cases, the system achieved an average benchmark score of 88.5/100, with scores ranging from 84 to 96. The results suggest that structured, multi-agent workflows can produce useful first-pass invention analyses across different technology domains.
Additionally, the benchmark illustrates an important distinction: evaluating legal AI is not simply a question of whether a model produces a plausible answer. The evaluation requires a defined task, a controlled test set, an explicit methodology, a reference standard, and a reproducible scoring framework.
Finally, future iterations could expand the test set to additional technical domains, including medical devices, robotics, security, video coding, autonomous systems, and other areas encountered in patent practice.