AI Systems
Harvey Tenet: Engineering AI for Long-Horizon Legal Work.
Understanding the model, the agent, the training environment, and the benchmarks behind Harvey Tenet.
About This Note
On August 20, Harvey announced its first post-trained open-weight model—Harvey Tenet.¹
The announcement is unusually technical. It discusses reinforcement learning, long-horizon agents, LoRA, training environments, specialized models, agent harnesses, and several different benchmarks.
If you are a legal professional rather than an AI researcher, it is easy to get lost.
This note takes a different approach—we start from zero. You do not need to know what an MoE model, LoRA, GSPO, RLM, or agent harness is. We introduce each concept only when we need it. Some of the terminology is technical because Harvey's legal AI system is technical. However, the underlying ideas are easier to understand than the terminology suggests.
Our goal is not to reproduce Harvey's technical announcement. The goal is to understand what Harvey actually wants to build, how Tenet was trained and evaluated, and what the approach may tell us about the future of legal AI.
A note on evidence. Throughout this note, we try to keep four things apart:
- what Harvey explicitly states in its announcement;
- what a primary source other than Harvey establishes;
- what is derived by calculation from published information; and
- what remains unknown.
What Problem Is Harvey Trying to Solve?
Assume you are a junior associate at a law firm. Consider two requests from a partner.
Request 1: What is the statute of limitations associated with a client matter?
Request 2: Review documents in a client matter, identify important legal issues, research any relevant material, and prepare a memo for the partner.
These are very different problems: the first is a question-answering problem; the second is a work-execution problem.
If you ask an AI system to carry out the work, it may easily provide the answer to the first request. The second request, however, requires the AI system to:
- understand the assignment;
- inspect a large collection of documents;
- decide what information matters;
- search for relevant material;
- use tools;
- keep track of what it has already reviewed;
- investigate multiple issues;
- connect findings across documents;
- determine when its investigation is sufficient; and
- produce a professional work product.
This is what Harvey calls "long-horizon, agentic legal tasks." There are two important words: "long-horizon" and "agentic." First, a long-horizon legal task requires an AI system to perform many individual actions before it can produce a final answer. Second, an agentic system is not designed simply to provide answers to a prompt. Instead, it is designed to take actions inside an environment.
So Harvey is not simply saying: "We built a better legal LLM called Tenet." A more useful interpretation of Harvey's announcement is this:
We can take an existing LLM, place it in a realistic professional-work environment, give it tools and tasks, train it against expert-defined outcomes, and build AI systems that can perform professional work over long sequences of actions.
In its simplest shape, that is the pattern this whole note is about:
- Model
- Agent
- Environment
- Work Product
That distinction is the key to understanding Tenet.
You Are a Poor Junior Associate!

Learning #1: A Legal AI System Is More Than an LLM
The first thing to understand is that Tenet is only one component of Harvey's legal AI system.
Based on publicly available information, and as illustrated in Figure 1, the system built around Tenet includes at least three layers:
- a model layer;
- an agents and harnesses layer; and
- an environment, data, and tool layer.
Note that Figure 1 is a reconstruction, not a Harvey architecture diagram. Harvey has not published a production architecture. The boxes are drawn from things Harvey states, and the dashed "Routing / Orchestration" connector in the middle is drawn dashed on purpose.
Model Layer
The model layer includes the underlying large language models:
- Tenet;
- an M&A Diligence model;
- a Review Table model; and
- a Firm Knowledge model.
Table 1. Large Language Models in Harvey's Legal AI System¹
| Capability | Based Model | Primary Problem |
|---|---|---|
| Tenet | Kimi K3 | General long-horizon legal work |
| M&A Diligence | GLM-5.2 | Large document/data-room investigation |
| Review Table | GLM-5.2 | High-volume structured extraction |
| Firm Knowledge | Qwen3.8-27B | Finding and using accumulated firm knowledge |
To build Tenet, Harvey started with Kimi K3, an open-weight model developed by Moonshot AI,² and post-trained it. The other three models were built by post-training as well, but from different starting points and by different methods.
There is no one model here that does everything. Harvey describes these four capabilities as separately trained, sitting on three different base models from three different developers.¹
Why four models instead of one? The most plausible reading is that the tasks genuinely differ—reviewing thousands of documents in an M&A deal is a different engineering problem from working a single client matter. Harvey does not explain the choice, so this is an interpretation rather than a disclosure, but the pattern is consistent across all four capabilities.
Agents and Harnesses Layer
What is an agent?
An agent is an AI subsystem that can take actions rather than simply produce one answer.
Instead of "Here is my answer," an agent might do this:
Search documents → open a document → inspect a provision → search for related provisions → save a finding → investigate another issue → revise its analysis → prepare a final report.
The agent therefore creates a trajectory, which is a sequence of actions and observations.
This becomes especially important when the task is large.
Harvey's agents and harnesses layer appears to include generalist agents, such as the long-horizon legal agent, as well as specialist agents corresponding to the specialized capabilities.
What is a harness?
A harness is the machinery that allows an agent to actually perform work.
Imagine two workplaces:
In Workplace A, you are told "Here are 500 files. Good luck."
In Workplace B, you are told "Here are 500 files." And you also have:
- search;
- a document reader;
- a filesystem;
- scripts;
- a workspace;
- tools;
- ability to organize files;
- ability to save work.
Those two workplaces give the same underlying model very different abilities.
The second is an agentic harness.
What is routing?
Harvey states that its specialized capabilities "can be deployed as tools and sub-agents that our model can route to in order to tackle specific tasks."¹
Unfortunately, Harvey does not name which model does the routing. The sentence states "our model," not "Tenet." Harvey does not describe the mechanism, does not say whether routing is automatic or user-selected, does not say whether any model was trained to route, and does not say whether any of this is running in production today.¹
Environment, Data, and Tool Layer
The third layer is the workplace itself. As illustrated in Figure 1, the environment, data, and tool layer includes the filesystem, documents, the firm knowledge corpus, retrieval and search, a bash shell and a code interpreter, and other tools.¹
This is the layer that turns Workplace A into Workplace B. It is also the layer that a law firm is most likely to already own in some form: the documents are the firm's, the knowledge corpus is the firm's, and the definition of what tools an agent should have is a professional judgment by the firm.
Figure 1. Harvey's Legal AI System (reconstructed from public information)

Learning #2: How Did Harvey Turn Kimi K3 into Tenet?
The basic story is surprisingly intuitive.
Harvey took an existing model—Kimi K3—and placed it into a simulated professional-work environment.¹
The environment includes things such as an assignment, documents, tools, a filesystem, a workspace, and an expert-defined description of what a good result should contain.¹
The model then attempts the task. Not once, but repeatedly.
A single rollout can exceed 1,000 turns and hundreds of thousands of tokens.¹ In other words, the model may take more than a thousand individual interaction and action steps while trying to complete one legal assignment.
What Is a Rollout?
A rollout is simply one complete attempt to solve a task.
Unlike conventional AI training, Harvey is evaluating the entire process of doing the work, not just whether one generated sentence resembles a reference answer.
Harvey ran this at considerable scale. Its appendix reports roughly 1,750 task environments, more than 10,000 rollouts per training epoch, and approximately 150 NVIDIA B300 GPUs running for two months.¹
Legal Task Becomes a Training Environment
Harvey's training environment includes three important things.
1. Partner instruction. For example: "Please review this matter and prepare a diligence memo." Harvey notes that these instructions average around 50 words and are deliberately written as a request from a partner rather than a detailed specification of the expected output.¹ That under-specification is a design choice, and an interesting one: issue-spotting is pushed into the matter rather than handed to the model in the prompt.
2. Client matter. A collection of documents and other materials, deliberately mixing key documents with peripheral ones, with the underlying legal issues distributed across several files.¹
3. Expert rubric. A detailed description of what the finished work should contain.
The rubric functions like a grading checklist. Harvey says that an average task includes around 50 binary pass/fail criteria, while the largest tasks can contain hundreds, and that each criterion is tied to a specific deliverable file.¹
So instead of merely asking "Did the AI give the correct answer?", the training system can ask questions such as:
Did it identify this issue? Did it review this document? Did it provide the required analysis? Did it support the conclusion? Did it complete the assignment? Did it do so efficiently?
This is a much closer representation of professional legal work.
What Does the Model Actually Learn?
Harvey does not necessarily need to teach Kimi K3 basic legal rules from scratch.
What Harvey actually taught Kimi K3 is the procedure for performing legal work.
For example, when doing diligence, the long-horizon agent may:
- systematically look for indemnification provisions;
- search for related provisions;
- trace the provisions to relevant documents;
- keep investigating until the matter is sufficiently covered; and
- produce a final report.
Thus, the training is intended to improve how the model uses its legal knowledge to execute a professional legal task.
That is one reason we find that Tenet is much more interesting than simply another "legal LLM."
LoRA: What Exactly Is Being Changed?
Harvey used LoRA, or Low-Rank Adaptation, during the Tenet training.¹
The simplest way to think about LoRA is this: instead of rewriting all of Kimi K3's original weights, add trainable components that modify how the existing model behaves.
The original Kimi K3 remains largely frozen. The LoRA components are what training updates.
At first glance, "low-rank" sounds like "small." Nevertheless, low-rank does not necessarily mean small. Kimi K3 is a very large Mixture-of-Experts (MoE) model, and that changes the arithmetic considerably.
A Short Detour: What Is an MoE Model?
In a conventional model, every parameter participates in processing every word. In a Mixture-of-Experts model, the neural network includes a very large number of small specialized sub-networks—"experts"—and a router selects only a few of them for each piece of text.
Let's make an analogy. A conventional model is one generalist who personally handles every question. An MoE model is a firm with hundreds of specialists and an intake process that routes each question to a handful of them. The firm is enormous; the number of people working on any single question is small.
According to Kimi K3's published configuration, it has roughly 2.8 trillion parameters in total, but only about 104 billion are active for any given token. It has 93 layers, of which 92 include experts, and each of those layers holds 896 routed experts, of which 16 are selected at a time.²
The ~500,000 Tensors, and Why the Number Is Interesting
Now the LoRA arithmetic may start to make sense.
Harvey states that it applied a rank-64 LoRA "over the full Kimi K3 network, adapting all attention, MLP, and routed-expert weights (~500k expert tensors)."¹
Where does a number like 500,000 come from? Using Kimi K3's published configuration:² 92 layers containing experts, 896 routed experts per layer, and three relevant weight matrices per expert:
92 × 896 × 3 = 247,296 expert weight matrices
247,296 × 2 = 494,592 LoRA tensors
That is roughly 500,000, which matches Harvey's stated figure closely. Note that this calculation is derived from Kimi K3's published configuration file.
The conceptual lesson is this: a model can be enormous while the portion of it that training actually touches stays comparatively narrow.
Reinforcement Learning: The Training Loop
Now we can understand the role of reinforcement learning (RL). Harvey's training loop looks roughly like the one illustrated in Figure 2 and the flowchart below¹:
- Legal Task
- Agent Attempts the Task
- Work Product
- Judge Evaluates
- Reward
- RL Update
Each pass through that loop leaves the model a little better at the desired behavior.
More importantly, as shown in Figure 2, the model does not make one attempt at a task. It makes a group of attempts at the same task. Note that Harvey's appendix describes eight rollouts per group, and eight groups per training step.¹ In particular, the attempts are scored against each other.
Who Is the Judge?
Harvey says it evaluated candidate judge models and used Kimi 2.6 as the training reward judge.¹ The judge reads the work product, evaluates it against the rubric, and produces the reward signal used during training—it is literally part of the learning loop.
That is a different job from grading the finished model on a benchmark, which we will discuss later, and it is worth keeping the two apart:
| Role | Model | What It Scores |
|---|---|---|
| Training judge | Kimi 2.6 | Each rollout, during RL training |
| Benchmark grader | Varies by benchmark | The finished model, on LAB / APEX Agents / RedlineBench |
GSPO: What Does It Do?
GSPO stands for Group Sequence Policy Optimization.
You do not need the math to follow the idea: the system generates multiple attempts, scores them, compares their outcomes, and uses those relative outcomes to determine how the model should be updated.
In simplified form: generate multiple trajectories → score them → calculate relative advantages → optimize the policy.
GSPO is the mathematical machinery that turns those rewards into an update to the model. It was published by the Qwen team at Alibaba in 2025, and one of its stated advantages is that it is more stable when training very large Mixture-of-Experts models, which is exactly what Kimi K3 is.³
The important point is not the acronym. It is that the model is learning from the quality of complete attempts to perform the work.
The Reward Is Not Just "Was It Right?"
Harvey describes the reward as a weighted combination of several things: the fraction of rubric criteria satisfied; a count of the underlying legal issues actually solved; a bonus for attempts that satisfy every criterion; and a small penalty term that favors more concise deliverables among attempts of comparable quality.
Figure 2 captures this as quality plus efficiency, and the efficiency half is genuinely unusual. Harvey made efficiency a training objective: the model was rewarded for reaching the same standard with fewer tokens. Harvey reports the practical result as cost staying roughly flat, at $5.92 per benchmark task against $5.62 for the base model,⁴ while the number of completed tasks nearly doubled.¹
The weights on those reward terms are not disclosed. That is a meaningful gap, because the balance between "satisfy more criteria" and "fully resolve issues" determines what kind of work the model learns to produce.
Figure 2. How Tenet Was Trained (reconstructed from public information)

Learning #3: How Is Tenet Evaluated?
After training, Harvey needs to answer a question:
Does the system actually work?
This is where benchmarks enter the picture: a standardized set of tasks used to compare systems. Figure 3 shows the same long-horizon agentic system trained in Figure 2 now producing work products, scored by three different benchmarks, each with its own grading protocol. There is no single "score." There is a score per benchmark, per metric, per grading protocol.
The Three Benchmarks
LAB is Harvey's own legal agent benchmark. It is designed around long-horizon legal work and uses detailed criteria to evaluate whether the agent completed the required work. It is open source, and it is large: more than 1,200 tasks across 24 practice areas, graded by more than 75,000 expert-written criteria.⁵
APEX Agents is an external professional-agent benchmark from Mercor. It evaluates agents performing long-horizon professional work in simulated workplaces. Harvey used its corporate-lawyer subset, whose tasks were written by practicing attorneys.⁶
RedlineBench is an external benchmark from Crosby, focused on multi-turn contract negotiation and redlining.⁷
Harvey reports that Tenet improves on both APEX Agents and RedlineBench, and states that the model had not seen these benchmarks during training.¹
The external benchmarks therefore provide the most interesting evidence in the announcement. Does behavior learned on Harvey's own training tasks help on someone else's task set, graded by someone else's judges? Harvey's answer is yes, and that is a meaningful claim, because a model that had merely memorized its own benchmark would show nothing of the kind.
Criteria Pass Rate Versus All-Pass: Why the Same Model Has Two Very Different Scores
LAB illustrates a very important problem in evaluating legal AI, and it is a problem lawyers will find familiar.
Suppose a legal assignment has 50 required criteria. An AI might satisfy 45 of 50.
That sounds like a 90% result. That number is the criteria pass rate.
But if the benchmark uses an all-pass metric, the task still fails—because all-pass means every required criterion must be satisfied.
| Metric | What It Asks | On a 50-Criteria Task |
|---|---|---|
| Criteria pass rate | What percentage of individual criteria were satisfied? | 45/50 satisfied → 90% |
| All-pass | Was every required criterion satisfied? | 45/50 satisfied → fail |
This is deliberately harsh. And there is a good reason for it.
Imagine a partner asks an associate: "Review this transaction and identify all material legal issues."
If the associate identifies 19 important issues but misses one critical issue, it may not be meaningful to say: "You got 95% of the assignment right." The missed issue may be the most important issue.
So all-pass is closer to asking: Did the AI successfully complete the matter?
than: How many individual answers did the AI get right?
Why All-Pass Produces Surprisingly Low Numbers
Harvey's research shows that an average LAB task has approximately 50 criteria.
Suppose that a model has a 90.8% criterion-level success rate, and suppose each criterion were independent. For 50 criteria: 0.908^50 ≈ 0.8%.
So a model could be right on roughly nine out of ten individual criteria and still have a very low probability of getting every single criterion right.
The actual observed all-pass rate is much higher than this hypothetical calculation—around 10.8% for Kimi K3, according to the public leaderboard maintained by Vals, which also reports that same model at a 90.8% criteria pass rate.⁸
That gap tells us something interesting: the criteria are not independent but correlated.
If an agent finds the right document, it may suddenly satisfy many related criteria. If it fails to find an important document, it may miss many related criteria.
So all-pass is partly measuring something like: How successfully did the agent cover the matter?
What Does Tenet's 19.7% Mean?
Fireworks AI, Harvey's training partner, reports Tenet at approximately 19.7% all-pass on LAB, compared with approximately 10.8% for Kimi K3.⁴
What does 19.7% mean?
It is important not to misunderstand it. 19.7% does not mean Tenet is 19.7% accurate.
It means approximately 19.7% of the evaluated matters satisfied all required criteria under the all-pass standard.
Conversely, about 80% of matters—roughly four in five—still had at least one unsatisfied criterion, something a reviewing attorney would need to catch.
Therefore, Tenet nearly doubles the number of matters achieving all-pass compared with Kimi K3—a large relative improvement on a genuinely hard metric. But it does not mean that Tenet can independently complete the majority of these legal assignments without review.
But What Caused the Improvement?
This is where the evaluation becomes particularly interesting, and where we want you to be most careful.
The publicly reported score for base Kimi K3 on Mercor's leaderboard is approximately 58.8%.⁹
Harvey then reports that same base model—untrained, unchanged—at 67.5% when operated in Harvey's own harness, which gives the model direct shell access to a filesystem containing the task's world rather than routing everything through Mercor's standard tool interface.¹
Tenet is then reported at 74.0%.⁴
So the headline improvement, 58.8 to 74.0, is not a single change. Two things differ between those endpoints: (1) the harness and (2) the trained model.
Harvey says this plainly, and also reports that the same harness substitution moved two other frontier models in opposite directions—one up by 7.7, one down by 4.2.¹ A harness that helps one model and hurts another is not a neutral measuring instrument.
Arithmetically, the two differences are about 8.7% and about 6.5%.
Now the important caveat, and we want to state it clearly because it is easy to over-read.
Those two numbers are not a clean decomposition. The first difference of 8.7% involves a different harness and a different organization running the evaluation, with run conditions that are not described. The second difference of 6.5% comes closer to isolating the model, since only the weights changed. Nevertheless, Tenet was trained in a working environment and then evaluated in one, so some part of that gain may be a model fitted to its evaluation setup rather than transferable capability.
One Step Further: A Benchmark Score May Belong to the System, Not Just the Trained Model
Suppose we say "Model X = 74%." That sounds like an intrinsic property of Model X. But the truth can be:
Model X + harness + tools + task set + grading protocol = 74%
Change the harness and the score can change. Change the metric and the score can change. Change the grader and the score can change.
The clearest illustration is one we have already seen: the same model, the same weights, scoring very differently depending on what surrounds it.
| Benchmark (Metric) | Harness / Reporting Source | Kimi K3 Score |
|---|---|---|
| LAB (all-pass) | Harvey's harness; reported by Fireworks | 10.8%⁴ |
| LAB (criteria pass rate) | Harvey's harness; reported by Vals | 90.8%⁸ |
| APEX Agents | Mercor's standard harness | 58.8%⁹ |
| APEX Agents | Harvey's own harness | 67.5%¹ |
This does not mean benchmarks are useless. It means benchmark results need more provenance.
For legal AI in particular, we would want to know at least:
Which model? Which benchmark? Which metric? Which harness? Which task set or split? Which grader? Which grading protocol?
A benchmark number without that information can be difficult to interpret—and, in a procurement conversation, easy to misread.
Figure 3. Benchmark Evaluation of Tenet (reconstructed from public information)

Learning #4: AI Infrastructure Is Part of the Story
Tenet did not emerge from one company working alone.
Harvey's acknowledgments name a number of collaborators: Fireworks, Engram, Baseten, Applied Compute, NVIDIA, Mercor, and Snorkel AI.¹ On a careful reading of the announcement, each appears to occupy a different position in the stack.
Fireworks AI is the training and inference infrastructure partner for Tenet itself. Harvey describes post-training the model "together with Fireworks research," and its appendix describes engineering that only a platform partner could supply—keeping thousands of long rollouts running while the trainer updates weights in place, and making the training system and the serving system produce numerically identical results so that the learning signal is not corrupted by the difference between them.¹ This is not a compute rental relationship. It is a research collaboration on infrastructure.
Mercor is the author of the APEX and APEX Agents benchmarks used to evaluate Tenet, and Harvey also credits it with helping build and scale the human expert datasets used in training and with reviewing synthetic data. Expert legal labor is an input to this system, and Harvey says so directly: expert data "made training at this scale possible."¹
The pattern is that modern AI development increasingly depends on an infrastructure stack:
compute → model serving → training infrastructure → agent framework → tools → data → expert evaluation → benchmarks.
The model is only one layer.
This is why AI infrastructure has become such an important area of the AI industry. And it explains something practical: a legal AI company does not need to build every component itself. Neither does a law firm.
Take-aways
1. Legal AI is moving beyond question answering. The harder problem is not "Can the AI answer a legal question?" It is "Can the AI execute a professional legal assignment?" That requires persistence, tool use, document management, reasoning, and quality control.
2. The model is only one component. Tenet is an LLM. But the system around Tenet includes agents, harnesses, tools, data, environments, and separately trained specialized models. We should increasingly think about legal AI systems rather than legal AI models.
3. The work environment can matter enormously. Harvey's own APEX comparison shows that changing the harness alone can materially move a model's score and that training a model to use a better environment mattered considerably more than providing the environment alone.
But how much of an AI system's professional capability comes from the model, and how much from the environment in which the model works?
We do not have a complete answer.
4. Professional knowledge and professional behavior are different problems. An AI may already know what a legal term means. That does not mean it knows how to conduct a complete diligence review. Training for professional work may be less about teaching isolated facts and more about teaching how to perform the work.
5. Evaluation needs to become more sophisticated. A single benchmark score can hide important differences. A legal AI system should be evaluated with attention to model, harness, task, metric, and grader. The more agentic the system becomes, the more this matters.
Conclusion
Harvey Tenet is interesting not simply because Harvey trained a new model.
It is interesting because it provides a glimpse into a broader direction for AI systems designed for professional work.
The basic pattern is:
Define the work. Build an environment where an AI can actually perform it. Define what good output looks like. Let the system attempt the work repeatedly, evaluate the results, use that evaluation to improve the system, and then test the improved system on tasks it never trained on.
This is fundamentally different from the traditional chatbot model—Prompt → answer. The emerging professional-AI model is closer to:
- Assignment
- Environment
- Agentic Work
- Work Product
- Evaluation
- Improvement
Tenet is one example of this approach in legal AI.
Defining the work, describing what good output contains, and deciding when a matter has been adequately covered are professional judgments. In Harvey's legal AI system, they are encoded in rubrics written by lawyers, and they are what the entire training apparatus is pointed at.
Therefore, the fundamental question for the legal AI field may not be:
"Who has the best legal LLM?"
but:
"Who can build the best system for getting professional work done?"
That is a question worth exploring, and it is a question in which the legal profession has more to contribute than it might assume.
Terminology
Agent. An AI subsystem that can take actions, use tools, and continue working toward a goal rather than simply produce one response.
Agentic system. A system in which an AI model operates through actions, tools, memory or workspace, and an execution environment.
All-pass. A benchmark metric in which a task counts as successful only if all required criteria are satisfied.
Benchmark. A standardized set of tasks used to evaluate an AI system.
Criteria pass rate. The proportion of individual criteria a system satisfies, averaged across tasks. Almost always much higher than the all-pass rate on the same run.
GSPO. Group Sequence Policy Optimization. The reinforcement learning optimization method Harvey used to update the Tenet policy based on groups of sampled trajectories. Published by Alibaba's Qwen team in 2025.³
Harness. The software and infrastructure surrounding a model that allows an agent to perform a task: its tools, workspace, filesystem, and the rules of its interaction with the environment.
Judge. A model or other evaluator that assesses the agent's work and produces a score or reward. Harvey's training judge and its benchmark graders are different models.
Long-horizon. A task requiring many sequential actions before completion.
LoRA. Low-Rank Adaptation. A method for adapting a model by training additional low-rank components while leaving most of the original model unchanged.
MoE. Mixture of Experts. A model architecture containing many expert sub-networks, with only a few activated for any given piece of text.
Post-training. Additional training applied to an already-trained model, to change its behavior rather than build it from scratch.
Reinforcement learning (RL). A training method in which the system uses rewards from evaluated behavior to improve the model.
Reward. A numerical signal representing how well the system performed according to the training objective.
Rollout. One complete attempt by the model or agent to perform a task.
Rubric. The expert-written checklist defining what a completed work product must contain. In Harvey's benchmark, criteria are binary and each is tied to a specific deliverable.
References
1. Calvin Qi, Vasudha Rengarajan, Julio Pereyra, Niko Grupen, and Gabe Pereyra (Harvey AI). "Harvey Tenet Research Preview." Harvey AI Blog, August 20, 2026. harvey.ai/blog/post-training-update-harvey-tenet
2. Moonshot AI. "Kimi-K3" (model card and configuration). Hugging Face. huggingface.co/moonshotai/Kimi-K3
3. Zheng, et al. (Qwen Team, Alibaba). "Group Sequence Policy Optimization." arXiv:2507.18071, July 2025. arxiv.org/abs/2507.18071
4. Fireworks AI. "Post-Training Kimi K3 with Harvey for Long-Horizon Legal Work." Fireworks AI Blog, August 26, 2026. fireworks.ai/blog/post-training-kimi-k3-with-harvey-for-long-horizon-legal-work
5. Niko Grupen, Gabe Pereyra, and Julio Pereyra (Harvey AI). "Introducing Harvey's Legal Agent Benchmark." Harvey AI Blog, May 6, 2026. harvey.ai/en-US/blog/introducing-harveys-legal-agent-benchmark
6. Mercor. "Introducing APEX-Agents." Mercor Blog, January 21, 2026. mercor.com/blog/introducing-apex-agents
7. Crosby Legal and micro1. "RedlineBench" (dataset). Hugging Face. huggingface.co/datasets/crosbylegal/RedlineBench
8. Vals AI. "Harvey Legal Agent Benchmark." Vals Blog, September 5, 2026. vals.ai/benchmarks/hlab
9. Mercor. "APEX-Agents Corporate-Lawyer Leaderboard." mercor.com/apex/apex-agents-leaderboard/corporate-lawyer-agent