Harvey LAB-AA Launches: First Agentic Legal Benchmark Across 24 Practice Areas
Harvey and Artificial Analysis jointly released Harvey LAB-AA (Legal Agent Benchmark), a new agentic legal benchmark evaluating language models on 120 real-world legal tasks across 24 practice areas — including corporate M&A, capital markets, tax, litigation, and bankruptcy. The primary metric is an "all-pass rate" requiring every criterion in a task rubric to be satisfied, reflecting the binary s
BY FRONTIER DESK · JULY 13, 2026 · 1 MIN READ
Harvey and Artificial Analysis jointly released Harvey LAB-AA (Legal Agent Benchmark), a new agentic legal benchmark evaluating language models on 120 real-world legal tasks across 24 practice areas — including corporate M&A, capital markets, tax, litigation, and bankruptcy. The primary metric is an "all-pass rate" requiring every criterion in a task rubric to be satisfied, reflecting the binary success standard of professional legal deliverables. At launch, Claude Fable 5 leads with a 14.2% all-pass rate at $18.9 per task; Claude Opus 4.8 and GLM-5.2 tie at 7.5%, with GLM-5.2 achieving that score at approximately 15% of Fable 5's cost ($1.3 per task). The best model still fails ~86% of professional legal deliverables on an all-pass basis — which means frontier legal AI is still far from autonomous. For operators and investors, the LAB-AA benchmark is more useful than BigLaw Bench for evaluating real production readiness: BigLaw Bench tests accuracy on discrete legal questions, while LAB-AA tests end-to-end completion of full legal work products, which is the actual commercial bar.