Can an AI agent actually reason like a lawyer? What a 162-task benchmark found
Regulatory watch agents are increasingly pitched as able to "read the law for you." A large benchmark built with legal professionals shows where that holds up — and where it does not.
A regulatory or legal watch agent is often pitched as something that can "read the law for you" — flag a new text, summarize a ruling, tell you whether a clause still holds. How far that claim actually goes is not a matter of opinion: a large interdisciplinary team of researchers and legal professionals built and published LegalBench, a benchmark of 162 tasks spanning several distinct families of legal reasoning, presented at NeurIPS 2023.
Retrieval is not reasoning
The benchmark’s tasks were hand-crafted by legal professionals precisely so that "legal reasoning" would not collapse into a single vague category. The pattern that comes out of this kind of structured evaluation, echoed across similar legal-AI benchmarks, is consistent: models perform far better on tasks that resemble information retrieval — finding the relevant rule, recalling a definition, locating the applicable text — than on tasks that require applying that rule to a specific, messy set of facts and reaching a defensible conclusion. Retrieval is a solved-enough problem; judgment is not.
What that means for a regulatory watch agent
The practical dividing line follows the same pattern. An agent is a strong fit for monitoring official sources, flagging when a text changes, and summarizing what changed in plain language — genuine retrieval-and-summarization work at a volume no person can sustain by manually checking dozens of registers every week. It is a poor substitute for the step after that: deciding what a given change means for your specific contract, your specific process, your specific liability. That step still needs a person who can be held accountable for the judgment call.
- →Delegate to the agent: monitoring official registers and gazettes, flagging new or amended texts, summarizing what changed since the last version.
- →Keep with a person: interpreting what a change means for your specific situation, and any decision or sign-off that follows from that interpretation.
- →Treat a confident-sounding legal summary the same way you would treat any other AI output on a consequential topic — as a first draft to verify, not a final answer.
Free resource
The self-assessment grid: 20 tasks AI can automate
Sales, admin, support, operations: the 20 tasks AI agents already handle in SMEs — with, for each one, the tell-tale sign that your team is concerned.
Read next
Business & compliance
EU AI Act: what became mandatory on August 2, 2026 — and what it means for your SMB
19 August 2026·5 min read
Business & compliance
Claude now watermarks its text: what Anthropic’s move changes — and doesn’t — for your SMB’s AI content
18 August 2026·5 min read
Business & compliance
Computer History: ChatGPT now remembers your activity — except in France (for now)
17 August 2026·5 min read