Legal Reasoning · AI Systems
Why Legal Reasoning Cannot Be Reduced to One AI Prompt
Prompt engineering matters, but difficult legal work also requires engineering the context, loops, and structures around language models.
The Problem
A complex legal task can involve many different steps. It may require understanding facts, identifying legal issues, finding and analyzing authorities, performing legal reasoning, and producing work products.
One common approach is to put all of those instructions into one large prompt. The idea is rather simple: tell an AI system everything it needs to do, provide it with relevant documents, and ask it to produce final work product.
Ironically, a prompt is not the same thing as a workflow.
A long prompt may tell an AI agent to perform ten or twenty different tasks of a legal workflow, but that does not mean the AI agent will reliably perform every one of them. It may skip a step, combine several steps, perform them in the wrong order, or produce a convincing final conclusion without completing all of the work that was requested. In other words, a long prompt can give an AI agent instructions, but there is no guarantee that the entire legal workflow was actually carried out.
This matters in legal work because the steps in a workflow are often connected. A mistake early in the analysis can affect everything that follows.
That raises a broader question: if prompting is part of the problem, what should we do?
The Approach
Here, we are thinking about legal AI systems as requiring several different kinds of engineering.
Prompt Engineering. It is the most familiar, which is about telling the model what to do: analyze an issue, compare several positions, draft a paragraph, review a document, or perform a certain task. Nonetheless, the prompt is only one part of the AI system.
Context Engineering. It is about deciding what a language model needs to see for performing a specific legal task. The relevant context may include legal documents, prior work products, attorney's instructions, internal know-how, factual findings, or legal authorities.
Loop Engineering. It is about what happens when one pass is not enough. A legal analysis may need to be reviewed, challenged, revised, and reviewed again. Instead of asking a language model to get everything right in one pass, the AI system can create a process around repeated analysis and review.
Graph Engineering. It is about how all pieces are connected. Which task comes first? Which tasks may appear at the same time? Which work product does a later stage depend on? Where should an attorney review the work? What happens when a stage fails?
In short, these four engineering ideas are intercorrelated:
- Prompt → What should a language model do?
- Context → What does a language model need to see?
- Loop → What happens when one pass is not enough?
- Graph → How are the AI agents, tasks, work products, and attorney reviews connected together?
Prompt engineering still matters. But as a legal task becomes more complex, the engineering problem moves beyond the prompt.
Figure 1. From Prompt Engineering to System Engineering

What We Are Investigating
We are investigating where different parts of a legal AI workflow should live.
For example:
- What should be handled by a prompt?
- What information should be provided through context?
- Which legal task should be separated into different stages?
- Which legal task needs a review or revision loop?
- Which intermediate result should become a persistent work product?
- Which stage can be reused in another legal workflow?
- What happens when an AI agent skips a step?
- What happens when an AI agent produces an incomplete result?
- When does an AI system become unnecessarily complicated?
Problems with a "Giant" Prompt
One particular problem we are interested in is the "giant prompt." A giant prompt may try to capture an entire legal workflow:
- Read the documents.
- Identify the issues.
- Analyze the relevant law.
- Find the authorities.
- Compare the arguments.
- Check the analysis.
- Review your own work.
- Revise the analysis.
- Draft the final report.
- ...
At first glance, this may look like a workflow. However, it is merely one set of attorney's instructions given to an AI agent. Still, the AI agent has to remember the instructions, decide how to execute them, keep track of what it has already done, and produce, hopefully, the expected outputs. As the number of instructions grows, it becomes harder to know whether the AI agent actually completed the intended process.
There is also another problem: portability.
A giant prompt designed for one type of legal workflow is often tightly connected to a specific legal task. If we want to use part of the same reasoning process in a different legal workflow, we may have to rewrite a large portion of the prompt.
Instead, a more structured AI system should be built from components that can be reused across different legal workflows. For example, a reasoning step or a review process should not have to be rewritten from scratch every time when the underlying legal task changes.
We are still investigating where that boundary should be.
Figure 2. “Giant” Prompt vs. Reusable Components

Why It Matters
The goal is not to eliminate prompts. Rather, the goal is to avoid making a prompt carry the entire legal workflow.
A complex legal task can be decomposed into smaller pieces, with each piece having its own instructions, context, work product, and review process. This can make an AI system easier to understand, and also make failures easier to locate.
If a single giant prompt produces a poor final answer, it may be difficult to determine whether the problem came from understanding certain facts, identifying relevant issues, analyzing legal authorities, or simply skipping some instructions. By contrast, if those tasks are separated into meaningful stages, the AI system then has more places where we can inspect what happened.
That is why we think about legal AI as a system, not simply as a collection of prompts. A language model is important, but the AI system around the model determines what the model is asked to do, what it sees, what happens after it responds, and how its work connects to everything else.
Conclusion
Prompt engineering is useful, but a complex legal task is usually more than a prompting problem. In particular, when too many instructions are placed into one large prompt, an AI agent may miss steps, combine tasks that should be separated, or produce a final conclusion without making the work between the beginning and the end visible.
Breaking a legal task into meaningful stages is only part of the solution. We also need to decide what a language model sees, what happens when one pass is not enough, and how the different stages connect. Prompt, context, loop, and graph engineering give us different ways to design those parts of the AI system.
That is what we are investigating: not simply how to write better prompts, but how to engineer the prompt, context, loops, and workflow around the language model, such that comprehensive legal work can move reliably from one stage to the next.