ooligo
ENTRY TYPE · framework

How to Read Legal AI Benchmarks

By Marius Bughiu Last updated 2026-08-21 Legal Ops

On the hardest public evidence available in August 2026, the leading model satisfies 94.6% of individual rubric criteria on real legal deliverables and produces a deliverable where every criterion passes on 14.2% of tasks. Both figures come from the same benchmark, the same harness, the same run. Neither is wrong. That 80-point spread is the entire content of a legal-AI accuracy claim, and it is why “how good is legal AI” has no single answer: the honest answer depends on whether a lawyer reads the output before it leaves the building. This page is the five-line scorecard for any legal-AI number you are shown, with the 2026 figures to calibrate against.

When to use this

Use it when a vendor, an analyst deck, or a leaderboard hands you a percentage and asks you to act on it — a pilot decision, a model choice inside a platform you already own, a budget line. Do not use it to interrogate a vendor’s claim about your contracts; that is a different procedure, and it lives in how to evaluate AI contract review accuracy claims. This page is about published benchmarks: what they measure, which number the publisher chose to lead with, and what the score does not survive contact with.

Line 1 — Which grading standard produced the number?

Every serious 2026 legal benchmark reports two metrics, and the difference between them is larger than the difference between any two vendors.

Criterion pass rate is the share of individual rubric criteria a model satisfies. All-pass rate is the share of tasks where every criterion passes, with no partial credit. Artificial Analysis publishes both for Harvey LAB-AA, its independent reimplementation of Harvey’s Legal Agent Benchmark, released 2026-07-07:

ModelCriterion pass rateAll-pass rateCost per task
Kimi K3 (max)94.6%
Claude Fable 5 (with fallback)93.6%14.2%~$18.90
Muse Spark 1.1 (xhigh)93.1%
Claude Opus 4.8 (max)7.5%~$8.20
GLM-5.2 (max)7.5%~$1.30

Harvey’s own run of the full benchmark reports the top model completing 7.1% of tasks end-to-end. Nothing here is in tension. A memo with 19 of 20 rubric criteria satisfied scores 95% on one metric and zero on the other, because the twentieth criterion was the limitation-of-liability position.

The calibration. If a reviewer reads the whole deliverable before it ships, criterion pass rate is the number that predicts how much editing they do. If the output goes to a client, a court, or a counterparty without a full read, all-pass is the only number that describes your exposure — and at 14.2%, no 2026 system supports that mode of use. A vendor quoting a mid-90s figure for an unreviewed workflow has answered a question you did not ask.

Line 2 — Short horizon or long horizon?

Benchmarks split cleanly by task length, and scores do not transfer across the split.

Short-horizon benchmarks grade one bounded judgment: LegalBench (162 hand-built tasks across six types of legal reasoning, NeurIPS 2023), CUAD’s 41 clause categories, LEXam, and Harvey’s earlier BigLaw Bench. Read a contract, answer a question, classify a clause.

Long-horizon benchmarks grade a work product. Harvey open-sourced LAB on 2026-05-06: more than 1,200 tasks across 24 practice areas, graded against more than 75,000 expert-written rubric criteria. Each task hands the agent a partner-style instruction and a matter’s documents, and requires it to plan, search, work across the materials, and produce a deliverable.

The calibration. Match the horizon to what you are buying. A clause-extraction score says nothing about an agent that drafts a memo, and a long-horizon all-pass rate understates a tool you only use for extraction. When a vendor selling an agentic product quotes a short-horizon score, the mismatch is the finding.

Line 3 — Whose harness, and whose judge?

A published number measures a model plus the scaffolding around it. Artificial Analysis runs LAB-AA on its own Stirrup agent harness and states the divergences plainly: it omits Harvey’s custom tools and document-generation scripts, uses simplified prompts, requires exact filename matches, and grades with a single Gemini 3.1 Pro judge rather than Harvey’s original method. Harvey’s own numbers on the same underlying tasks are not the same numbers.

The calibration. LAB-AA scores a model. A vendor’s own score covers a model plus retrieval, tools, prompts, and guardrails they built. Those are not comparable quantities, and the gap between them is exactly the product. Ask which one you are looking at, and refuse the comparison when a vendor puts its product score in a table next to raw model scores.

Line 4 — Is the task set private and held out?

LAB-AA runs on 120 private tasks; Harvey published the code and a portion of the dataset on GitHub. That split is deliberate. A public task set is a training target — once the questions are on the internet, a high score stops distinguishing a model that reasons from one that has read the answers.

The calibration. A private held-out set run by a third party is the strongest evidence tier. A public set run by a third party is the middle tier — usable for ranking, weak for absolute claims. A public set run by the vendor is a marketing artifact. Ask when the set was last rotated; a benchmark that has not changed in two years is measuring memorisation.

Line 5 — What did it cost per task?

Cost per task is published alongside quality on LAB-AA, and dividing one by the other produces the number that drives a build-or-buy decision: cost per accepted deliverable = cost per task ÷ all-pass rate.

ModelCost/taskAll-passCost per accepted deliverable
GLM-5.2 (max)$1.307.5%~$17
Claude Opus 4.8 (max)$8.207.5%~$109
Claude Fable 5 (with fallback)$18.9014.2%~$133

The 14.5× spread in raw model cost compresses to about 8× once quality is priced in — and then stops mattering. Assume a reviewer billing $400/hour, which is an assumption rather than a measured figure: $133 buys 20 minutes of that person’s attention. The entire model-cost range across the leaderboard is worth 3 to 20 minutes of one lawyer’s time per deliverable.

The calibration. Below roughly 30 minutes of reviewer time per deliverable in model spend, pick on all-pass rate and ignore price. Above it, model cost has become a real line item and belongs in the business case. Most 2026 legal workloads sit well below that threshold, which means teams optimising legal AI spend on token price are optimising the wrong term. Seat-based vs usage-based AI pricing covers how vendors package that cost once it reaches your invoice.

The human baseline is a moving comparison, not a constant

Vals AI’s legal research report, published 2025-10-14, scored 200 questions across 10 question types (one disregarded for a formulation error) on a weighting of accuracy 50%, authoritativeness 40%, appropriateness 10%. Every AI product landed within four points of the others, at 74% to 78%, against a lawyer baseline of 69%. Multi-jurisdictional questions cost every system an average of 11 points.

Two things travel with that result. Thomson Reuters and LexisNexis declined to participate and vLex withdrew, so the three largest legal research platforms are absent from the strongest independent research benchmark on the market — an unmeasured vendor is not a losing vendor, and it is not a winning one either. And direction reverses by task: the February 2025 Vals study found tools beating the lawyer control group on document extraction while losing to it badly on redlining, using the same vendors and the same panel.

Watch-outs, each with its guard

A single headline percentage. One number means the publisher chose which of the two metrics to show. Guard: refuse to act on any legal-AI score until you have both the criterion pass rate and the all-pass rate, or the publisher’s statement of which one it is.

Comparing a product score to a model score. Vendor scoreboards place their platform beside raw frontier models, where the harness difference does the work. Guard: require both rows to come from one harness, or read them as separate tables.

Assuming benchmark rank predicts your matter type. LAB-AA spans 24 practice areas and reports an average; the variance across areas is not published per model. Guard: run 20 of your own real matters through the shortlist before the contract, and score them with your own rubric — the grounding and hallucination checks belong in that same pass.

Treating a score as a shelf life. Models on this leaderboard turn over in weeks, and 14 of 37 listed models carried a LAB-AA result at the time of writing. Guard: re-check the leaderboard before renewal, and write model-substitution rights into the order form so a platform can move you to a better model without a new negotiation.

When this framework breaks down

It breaks on non-English and non-US work, where none of the public benchmarks have coverage worth acting on, and on firm-specific work product where the rubric that matters is your own precedent bank rather than an expert panel’s. It also breaks where the deliverable has no answer key — strategy memos, negotiation posture, client counselling. There the benchmark tells you which model is a better starting point and nothing else, and the evaluation has to be a supervised pilot on your own matters.

For where these scores land in a buying decision, read legal AI vs legaltech and vertical AI model vs GPT wrapper, then the vendor pages for Harvey, CoCounsel, Legora, Lexis+ with Protégé, and vLex Vincent AI.