Test 003 · Attorney Judgment
Attorney Feedback and Revision Loop
Testing whether attorney judgment can be reliably translated into AI revisions.
Attorney Review · Revision · Human-in-the-Loop · Work Products
Aug 2026
01 · Question
Can attorney judgment be reliably translated into AI revisions?
In ordinary legal practice, a senior attorney's instruction to a junior attorney relies on shared context: the junior attorney can draw on professional judgment, on familiarity with the matter, and on the simple option of asking a follow-up question when the instruction is unclear. Nevertheless, a legal AI workflow does not have access to that shared context. An attorney's instruction (i.e., attorney feedback) must instead be interpreted by a computational system, which cannot safely fill in missing intent the way a human attorney might infer it.
This test examines how attorney feedback enters a structured legal AI workflow, and what the feedback needs to include for the system to correctly operationalize attorney judgment. In particular, the underlying question is whether attorney judgment can be reliably translated from natural language into a structured revision plan that an AI agent can actually execute in a revision loop.
Our hypothesis is that the reliability of the translation depends on whether attorney feedback includes sufficient information to identify the intended change. Specific and actionable feedback should produce more reliable revision actions, while ambiguous, underspecified, or conflicting feedback should increase the need for clarification.
02 · Method
Figure 1 · Feedback Interpretation
How an AI agent responds to different levels of instruction clarity.

System
The system under test is the structured, multi-agent Invention Analysis workflow described in We Build → Workflows.
Here, the revision loop does not pass attorney feedback directly to the Disclosure Analysis Agent (i.e., Disclosure Analyzer) that produces a revised Invention Analysis Report. Instead, a dedicated Feedback Review Agent (i.e., Feedback Reviewer) reviews the attorney feedback first and produces a Revision Plan: an explicit record of which feedback items are accepted, rejected, or require clarification, and which Revision Actions each accepted item authorizes.
This intermediate step separates interpreting the attorney's intent from executing a change. The Disclosure Analyzer responsible for revising the work product does not reinterpret attorney feedback or resolve conflicts on its own — it applies the Revision Actions that the Revision Plan already specifies. Making the intended changes explicit before revision begins is what allows a revision to be inspected, and where necessary contested, rather than trusted solely on the basis of the final work product.
Conceptual architecture of the revision loop:
- Attorney Feedback
- Feedback Review Agent
- Revision Plan
- Revision
- Revised Work Product
Controlled Variables
For each revision cycle, the following remained conceptually constant:
- Legal task: invention analysis
- Workflow: structured Invention Analysis workflow
- Feedback mechanism: structured Attorney Feedback file
- Revision architecture: Feedback Reviewer → Revision Plan → Revision
- Work-product structure: revised Invention Analysis Report
- Evaluation focus: whether attorney instructions were correctly interpreted and translated into Revision Actions
The test corpus is not a controlled experiment in which feedback specificity was deliberately randomized. It is an accumulated set of revision cycles that, together, span a range of feedback forms:
- approval / no-change feedback;
- specific, actionable revision requests;
- multi-part revision instructions;
- structural changes to a conceptual hierarchy;
- ambiguous or underspecified instructions; and
- conflicting instructions.
This test examines how the workflow handled each form, not whether one form is statistically superior to another.
03 · Results & Discussion
Results
| Feedback Pattern | Observed System Behavior | Result |
|---|---|---|
| Approval / no-change feedback | The Feedback Reviewer recorded that no revision action was required. | The revision loop was complete without altering the work product. |
| Specific revision instructions | A feedback item naming a specific target was translated into a targeted Revision Action. | The requested change was applied and reflected in the revised work product. |
| Multi-part instructions | Feedback introducing a new element was translated into multiple related Revision Actions. | Downstream sections were updated to remain consistent with the new element. |
| Structural instructions | A request to restructure a conceptual hierarchy required identifying which dependent items needed to be remapped. | The hierarchy was restructured and dependent items were remapped accordingly. |
| Ambiguous instructions | Feedback that did not clearly identify the intended change gave the Feedback Reviewer insufficient basis to generate a Revision Action. | The item was flagged for clarification rather than acted on, since no reliable revision could be inferred from the feedback alone. |
| Conflicting instructions | Feedback items whose combined effect could not be simultaneously satisfied were flagged during consistency review. | Revision was withheld and attorney clarification was requested rather than resolving the conflict unilaterally. |
Discussion
The revision loop's behavior across the test corpus follows a rough progression of difficulty.
For a straightforward instruction naming a single, clearly identified change, the Feedback Reviewer translates attorney intent into a Revision Action with little ambiguity. For example, one revision cycle included the instruction:
"Rename Inventive Concept 2 (IC2) to more accurately reflect the underlying technical approach."
The workflow then incorporated the renamed inventive concept throughout the work product, updating each place that inventive concept was referenced rather than only its initial definition. This suggests the revision process tracks an inventive concept's identity across the report rather than treating each reference as an independent string to match.
More complex instructions require the Feedback Reviewer to do more than translate a single sentence into a single action. When feedback introduces a new element — for example, an additional technical feature the attorney wants reflected in the analysis — the Feedback Reviewer shall identify not only the immediate addition but its downstream consequences:
"Add a technical feature for the fallback recovery mechanism, and update the classification and patentability analysis accordingly."
The revision record for this instruction shows the classification and patentability sections being updated alongside the primary addition, rather than the new technical feature being inserted in isolation.
Moreover, structural instructions raise a different kind of difficulty. Consider an instruction to combine two Inventive Concepts:
"Combine Inventive Concept 2 (IC2) and Inventive Concept 3 (IC3)."
Combining two inventive concepts does not simply add or modify a single item — it changes the conceptual hierarchy the rest of the report depends on. Technical features tied to either original inventive concept have to be remapped to the combined inventive concept rather than left pointing at an inventive concept that no longer exists in its original form. Executing this kind of instruction requires the revision process to preserve internal consistency across the report, not just apply an isolated edit.
Additionally, conflicting instructions present a fundamentally different problem. In one recorded instance, one instruction directed the system to delete an inventive concept, while another instruction required a distinguishing feature to continue referring to that inventive concept:
"Delete Inventive Concept 2 (IC2)." / "Distinguishing Feature 3 (DF3) should continue to refer to IC2."
These two instructions cannot both be satisfied as written. The Feedback Reviewer detected that they could not be jointly satisfied and did not attempt to resolve the conflict on its own. This indicates that the Feedback Reviewer treats conflicting instructions as a category distinct from underspecified ones: not simply "not enough information," but "the information that exists is inconsistent." There is no reliable basis for inferring which instruction reflects the attorney's actual intent — whether the reference should be updated to point elsewhere, or whether the deletion should not be applied as written. In this situation, stopping and requesting clarification is preferable to silently pursuing an interpretation.
In short, this progression points to a broader observation about the role of attorney judgment in a workflow. Human judgment is not simply inserted at the end of a workflow. It becomes an input to the workflow, and therefore must itself be interpreted, structured, and incorporated in a form that downstream AI agents can reliably act upon. Attorney feedback that is specific and actionable can be operationalized directly. Attorney feedback that is ambiguous or self-contradictory cannot be operationalized without more information.
04 · Limitations
This test is an engineering evaluation of one structured legal AI workflow's revision loop, not a general study of attorney-AI interaction:
- The test corpus includes heterogeneous revision scenarios accumulated over the course of development, not a randomized controlled experiment.
- The observations concern whether attorney feedback was operationalized within this specific workflow, and should not be generalized to all legal tasks or all AI systems.
- Attorney feedback quality is itself difficult to define objectively; "specific" and "ambiguous" feedback were not assigned by a fixed, independent rubric.
- A successful revision — one that correctly applies the requested Revision Actions — does not establish that the resulting legal analysis is substantively correct.
- The test does not establish that any particular feedback style is universally optimal, only that the workflow's reliability varied with the information the feedback contained.
05 · Conclusion
The test extends a pattern observed elsewhere: reliability in a legal AI workflow depends on more than the capability of the underlying models and AI system architectures. Attorney feedback also needs to be translated into a structured interface before it can reliably drive downstream AI work.
A feedback reviewing agent exists because attorney feedback, in its natural-language form, is not directly actionable by an analyzing/drafting agent that performs a revision. Attorney feedback shall be reviewed, checked for consistency, and converted into an explicit revision plan before any change is made — and when that feedback cannot be reconciled into a single, coherent set of instructions, the workflow is designed to stop and ask rather than guess.
The broader lesson is that AI systems need interfaces not only between software components and AI agents, but also between attorney judgment and the computational workflow that acts on it. That is not a general theorem about AI and attorney oversight — it is what this particular workflow needed in order to treat attorney feedback as something more reliable than a suggestion.