Harvey LAB-AA Exposes AI Legal Agents Passing Under 10% of Real Tasks
Artificial Analysis and Harvey add a hallucination gate to their legal agent benchmark, collapsing top model pass rates from 27% to under 10%.
- Harvey LAB-AA v1.1 adds a hallucination gate that zeros any task containing a material hallucination.
- Grok 4.7 (xhigh) leads at 9.4% Hallucination-Gated All-Pass Rate, ahead of Muse Spark 1.3 at 8.9%.
- Over 60% of otherwise passing results contain a material hallucination across tested models.
- GPT-6 Astra averages 0.03 material hallucinations per task; Gemini 3.8 Flash averages 13.96.
- Rubric grading now uses a three-judge panel (GPT-6 Sol, Grok 4.7, Claude Opus 5.5) replacing v1.0's single judge.
- Top models aren't the most expensive: Grok at $9.50/task beats Claude Fable at $21.70/task.
Hallucination gate pushes legal-agent pass rates below 10%
Version 1.1 of the Harvey LAB-AA benchmark has reshuffled the leaderboard for frontier models performing legal work. A task now counts as a pass only when the deliverable satisfies every rubric criterion and contains no material hallucinations. Under that rule, the leading model passes 9.4% of tasks.
Harvey LAB-AA places agents in a sandbox with case documents and partner-style instructions. They must produce deliverables such as M&A change-of-control reports, deposition outlines, and disclosure schedules. The benchmark uses 120 private tasks across 24 legal practice areas, with task-specific rubrics grading each requirement.
One false claim can erase a pass
The revised methodology separates three measures that capture different kinds of performance:
- Criterion Pass Rate: the percentage of individual rubric criteria the model satisfies.
- All-Pass Rate: the percentage of tasks on which the model satisfies every rubric criterion.
- Hallucination-Gated All-Pass Rate: the percentage of tasks that satisfy every criterion and contain no material hallucinations.
A material hallucination could mislead a reader on a substantive issue, such as stating the wrong contractually required date. One such error disqualifies the entire task under the gated metric. Minor hallucinations, which are unlikely to change the legal interpretation, are reported separately.
Across the tested models, more than 60% of results that otherwise passed contained at least one material hallucination. The gate produced the following leaderboard:
| Model | All-Pass Rate | Gated Rate | Approx. Cost per Task |
|---|---|---|---|
| Grok 4.7 (xhigh) | Not reported | 9.4% | $9.50 |
| Muse Spark 1.3 (max) | 26.7% | 8.9% | $4.20 |
| GPT-6 Astra (max) | 8.9% | 8.6% | Not reported |
| GPT-6.1 Sol (max) | 7.5% | 6.9% | Not reported |
| Kimi K3 (max) | 16.7% | 5.3% | Not reported |
| GLM-5.3 (max) | 13.9% | 0.3% | Not reported |
Muse Spark 1.3 would have led the original All-Pass ranking at 26.7%, but material hallucinations reduced its gated score to 8.9%. GLM-5.3 fell from 13.9% to 0.3%. GPT-6 Astra moved from joint tenth to third because its score changed only slightly, from 8.9% to 8.6%.
Grounding scrambles the rankings
Criterion coverage and factual grounding diverge sharply in the results. Kimi K3 satisfies 93.0% of individual criteria while averaging 2.09 material hallucinations per task. Muse Spark 1.3 reaches a 96.0% Criterion Pass Rate but averages 1.68 material hallucinations.
GPT-6 Astra averages 0.03 material hallucinations per task, with four across all 120 tasks. Gemini 3.8 Flash averages 13.96 per task. These differences explain why a model can perform well on rubric coverage and still lose most of its task-level passes under the hallucination gate.
Cost also provides little guidance about gated performance. Grok 4.7 leads at about $9.50 per task, less than half the roughly $21.70 cost of Claude Fable 5.1 with maximum effort and fallback enabled. GPT-6 Luna costs about $0.22 per task and scores 3.3%, placing it on the Pareto frontier, where no tested option is both cheaper and higher-scoring.
Two passes check every claim
Hallucination review uses GPT-6 Sol at high effort in a two-stage pipeline. The first pass compares each deliverable with the supplied documents and flags three categories of claim:
- Statements that contradict the source material
- Content presented as sourced but absent from the record
- Specific assertions with no support in the provided documents
The second pass checks each flag against the source files, dismisses unsupported flags, and classifies the remaining errors as material or minor. This recheck reduces the effect of false positives from the initial review.
General legal knowledge falls outside the hallucination check when it is absent from the supplied documents. A correct statement of case law or statute is therefore not flagged solely because the task packet does not contain it.
The checker shapes the score
The benchmark team compared six checker models on a 20-task subset before choosing GPT-6 Sol for production. Their thresholds differed substantially: GPT-6 Sol upheld 470 material hallucinations, Claude Opus 5.5 upheld 99, and Claude Sonnet 5.5 upheld 57.
Rubric grading now uses a three-model panel comprising GPT-6 Sol, Grok 4.7, and Claude Opus 5.5. The benchmark averages their results to reduce self-preference bias, replacing the single judge used in version 1.0.
Checker disagreement remains a constraint on interpretation because the hallucination gate relies on one production model after the second-pass review. The private task set also prevents outsiders from reproducing task-level grading from the published methodology alone. Reported scores therefore describe performance under this dataset, agent setup, and judging stack.
What agent builders should measure
Harvey’s preference study found that legal experts treat hallucinations as a primary factor when choosing between otherwise comprehensive answers. Evaluation systems that reward completeness without testing claims against source material can overstate an agent’s readiness for professional workflows.
- Gate complete tasks: require every mandatory criterion and zero material unsupported claims.
- Keep metrics separate: report criterion coverage, task completion, material hallucinations, and minor hallucinations independently.
- Trace factual assertions: connect generated claims to source passages so validators and reviewers can inspect their support.
- Test judge sensitivity: compare automated checkers with human review and monitor whether rankings change with the evaluator.
- Track efficiency: measure cost and token use alongside grounded task success.
Token use shows why efficiency belongs in the evaluation. GPT-6 Astra produces about 81,000 output tokens per task and scores 8.6%, while Grok 4.7 uses roughly 180,000. The three Claude models generate between 202,000 and 562,000 output tokens per task yet score between 2.8% and 6.4%. Longer runs may create more opportunities for unsupported assertions without improving the final legal deliverable.
The benchmark’s 9.4% leading score applies to a controlled sandbox with private tasks and model-based judges. Production legal agents still require source tracing, validation, human review, and monitoring tailored to the documents and risks of each workflow.