Test 002 · Workflow Performance
Workflow Reliability and LLMs
Testing whether a structured legal workflow remains reliable when its underlying LLM changes.
Workflow Execution · Reliability · LLMs · Multi-Agent AI
Aug 2026
01 · Question
Can a structured legal workflow remain reliable when its underlying LLM changes?
Figure 1 · Workflow Interface
JSON defines the contract between workflow stages.

A structured AI workflow often passes the output of one stage to the next. For example, one AI agent may analyze a legal document and produce a structured work product that another agent uses as the basis for its analysis.
To make this handoff reliable, the work product needs a predictable format. In our workflows, some intermediate work products are represented as JSON (JavaScript Object Notation), which is a machine-readable format that organizes information into defined fields. This allows the next stage of the workflow to read and process the output automatically rather than relying on a human to interpret it.
Nevertheless, this architecture creates a practical failure point: an LLM can produce useful legal analysis but still produce outputs that the next stage cannot process. In particular, the JSON produced by the LLM may be malformed, or it may omit or misformat fields that the workflow expects.
This test examines what happens to these structured handoffs when the underlying LLM is changed. Our hypothesis is that changing the underlying LLM may change the reliability of these structured handoffs, even when the workflow, input, and output requirements remain the same.
02 · Method
Test Configuration
Table 1 · Test Configuration
| Field | Value |
|---|---|
| Input | Invention disclosure (fixed) |
| Workflow | Invention Analysis workflow (fixed) |
| LLMs | LLM A vs. LLM B |
| Runs | 10 executions per LLM |
| Output contract | Fixed JSON schema |
LLM A and LLM B represent two LLM configurations evaluated under the same workflow, input, prompts, and output contract. The labels are used to focus the test on the behavior of the workflow rather than on model branding. The observed differences in execution were not independently attributed to model capability; transient service conditions or other runtime factors may also contribute to the observed failures.
System
The system under test is the structured, multi-agent Invention Analysis workflow described in We Build → Workflows.
Controlled Variables
Ten executions were performed for each LLM to observe whether workflow failures occurred repeatedly under otherwise equivalent conditions.
Kept Constant
- Invention disclosure
- Workflow definition
- Agent/task definitions
- Prompts
- JSON schema
- Orchestration logic
- Downstream processing
Changed
- Underlying LLM
Evaluation Method
Each execution is evaluated at multiple checkpoints in the workflow rather than only on final completion.
- Output validity. Is the generated output valid JSON?
- Schema compliance. Does it conform to the expected structure?
- Stage ingestion. Can the next workflow stage consume it?
- Workflow execution. Does the workflow proceed without manual intervention?
- Completion. Does the pipeline generate the expected final work products?
Success rate = executions completed at a checkpoint / total executions
03 · Results & Discussion
Figure 2 · Experimental Comparison
Changing the LLM while other conditions remain constant.

Results
Table 2 · Workflow Execution Success Rates
| Measure | LLM A | LLM B |
|---|---|---|
| Valid JSON | 10/10 | 10/10 |
| Schema-compliant | 10/10 | 5/10 |
| Stage ingestion | 10/10 | 5/10 |
| Workflow execution | 10/10 | 5/10 |
| Pipeline completion | 10/10 | 5/10 |
Experimental scope. These results reflect executions conducted during a defined test period and under the service conditions observed at that time. They do not establish that the same execution rates would hold continuously or uniformly across all times of day or operating conditions.
Discussion
The observed failures illustrate a distinction between LLM capability and workflow reliability. In particular, an LLM may produce substantively useful output while still failing to satisfy the structural contract required by downstream components.
- LLM Output
- Structured Work Product
- Validation
- Downstream Stage
- Workflow Completion
This means that changing an LLM may change not only the quality of AI response, but the behavior of the system built around that response.
LLM selection is therefore a workflow engineering decision, not merely a model-quality decision.
A reliable workflow must treat structured outputs as interfaces that require validation, error detection, and recovery rather than assuming that an LLM will always produce machine-consumable output.
04 · Limitations
This test is an exploratory reliability experiment rather than a statistical characterization of LLM reliability:
- The test uses a limited number of executions.
- Testing is performed through external LLM APIs whose service conditions may vary over time.
- Observed failures may therefore reflect LLM behavior, API/service conditions, or their interaction.
- The test uses a single workflow and thus does not establish reliability across all legal workflows.
- The test focuses on workflow execution and structured outputs rather than the substantive legal reasoning.
05 · Conclusion
This test illustrates that workflow reliability can depend on the behavior of the underlying LLM, even when the workflow definition and input remain unchanged.
Structured outputs create an interface between probabilistic AI components and deterministic workflow software. Failures at that interface can interrupt downstream execution even when the underlying task has otherwise been performed successfully.
Reliable legal AI systems therefore require more than capable LLMs. They require explicit output contracts, validation, failure detection, and recovery mechanisms around those LLMs.