Skip to content
Legal AI Laboratory
← We Test

Test 002 · Workflow Performance

Workflow Reliability and LLMs

Testing whether a structured legal workflow remains reliable when its underlying LLM changes.

Workflow Execution · Reliability · LLMs · Multi-Agent AI

Aug 2026

01 · Question

Can a structured legal workflow remain reliable when its underlying LLM changes?

Figure 1 · Workflow Interface

JSON defines the contract between workflow stages.

Diagram showing JSON as the interface/contract between two workflow stages: an agent produces a JSON structured work product defined by a schema, required fields, and data types; if the output violates that contract the next stage cannot reliably process it, illustrating that substantive quality is not the same as contract compliance.

A structured AI workflow often passes the output of one stage to the next. For example, one AI agent may analyze a legal document and produce a structured work product that another agent uses as the basis for its analysis.

To make this handoff reliable, the work product needs a predictable format. In our workflows, some intermediate work products are represented as JSON (JavaScript Object Notation), which is a machine-readable format that organizes information into defined fields. This allows the next stage of the workflow to read and process the output automatically rather than relying on a human to interpret it.

Nevertheless, this architecture creates a practical failure point: an LLM can produce useful legal analysis but still produce outputs that the next stage cannot process. In particular, the JSON produced by the LLM may be malformed, or it may omit or misformat fields that the workflow expects.

This test examines what happens to these structured handoffs when the underlying LLM is changed. Our hypothesis is that changing the underlying LLM may change the reliability of these structured handoffs, even when the workflow, input, and output requirements remain the same.

02 · Method

Test Configuration

Table 1 · Test Configuration

FieldValue
InputInvention disclosure (fixed)
WorkflowInvention Analysis workflow (fixed)
LLMsLLM A vs. LLM B
Runs10 executions per LLM
Output contractFixed JSON schema

LLM A and LLM B represent two LLM configurations evaluated under the same workflow, input, prompts, and output contract. The labels are used to focus the test on the behavior of the workflow rather than on model branding. The observed differences in execution were not independently attributed to model capability; transient service conditions or other runtime factors may also contribute to the observed failures.

System

The system under test is the structured, multi-agent Invention Analysis workflow described in We Build → Workflows.

Controlled Variables

Ten executions were performed for each LLM to observe whether workflow failures occurred repeatedly under otherwise equivalent conditions.

Kept Constant

  • Invention disclosure
  • Workflow definition
  • Agent/task definitions
  • Prompts
  • JSON schema
  • Orchestration logic
  • Downstream processing

Changed

  • Underlying LLM

Evaluation Method

Each execution is evaluated at multiple checkpoints in the workflow rather than only on final completion.

  1. Output validity. Is the generated output valid JSON?
  2. Schema compliance. Does it conform to the expected structure?
  3. Stage ingestion. Can the next workflow stage consume it?
  4. Workflow execution. Does the workflow proceed without manual intervention?
  5. Completion. Does the pipeline generate the expected final work products?

Success rate = executions completed at a checkpoint / total executions

03 · Results & Discussion

Figure 2 · Experimental Comparison

Changing the LLM while other conditions remain constant.

Side-by-side comparison of LLM A and LLM B: LLM A's structured output is valid, compatible JSON that passes the parser/validator and lets the pipeline continue; LLM B's output is invalid or incompatible JSON that fails validation and interrupts the pipeline or requires intervention.

Results

Table 2 · Workflow Execution Success Rates

MeasureLLM ALLM B
Valid JSON10/1010/10
Schema-compliant10/105/10
Stage ingestion10/105/10
Workflow execution10/105/10
Pipeline completion10/105/10

Experimental scope. These results reflect executions conducted during a defined test period and under the service conditions observed at that time. They do not establish that the same execution rates would hold continuously or uniformly across all times of day or operating conditions.

Discussion

The observed failures illustrate a distinction between LLM capability and workflow reliability. In particular, an LLM may produce substantively useful output while still failing to satisfy the structural contract required by downstream components.


  1. LLM Output
  2. Structured Work Product
  3. Validation
  4. Downstream Stage
  5. Workflow Completion

This means that changing an LLM may change not only the quality of AI response, but the behavior of the system built around that response.

LLM selection is therefore a workflow engineering decision, not merely a model-quality decision.

A reliable workflow must treat structured outputs as interfaces that require validation, error detection, and recovery rather than assuming that an LLM will always produce machine-consumable output.

04 · Limitations

This test is an exploratory reliability experiment rather than a statistical characterization of LLM reliability:

  • The test uses a limited number of executions.
  • Testing is performed through external LLM APIs whose service conditions may vary over time.
  • Observed failures may therefore reflect LLM behavior, API/service conditions, or their interaction.
  • The test uses a single workflow and thus does not establish reliability across all legal workflows.
  • The test focuses on workflow execution and structured outputs rather than the substantive legal reasoning.

05 · Conclusion

This test illustrates that workflow reliability can depend on the behavior of the underlying LLM, even when the workflow definition and input remain unchanged.

Structured outputs create an interface between probabilistic AI components and deterministic workflow software. Failures at that interface can interrupt downstream execution even when the underlying task has otherwise been performed successfully.

Reliable legal AI systems therefore require more than capable LLMs. They require explicit output contracts, validation, failure detection, and recovery mechanisms around those LLMs.

Related

Related work will be linked here as it is published.