✒️ Blog / Practice notes

The Honest Limits of AI Contract Review: What Still Needs a Human Eye

AI contract review is genuinely good — and its failures are structural and predictable. Here's how it goes wrong, what the evidence says, and the checklist of what still belongs to you.

Every article in this series has ended at the same place: the machine finds, the lawyer decides. That is not a slogan — it is a description of where the technology reliably ends and where the profession reliably begins. This article is the honest companion piece: a specific accounting of how AI contract review fails, what the measured error rates actually are, and the concrete checklist of what still needs a human eye.

The evidence first

The most careful independent benchmark is the Stanford study led by Magesh, Dahl, and Ho — first released in 2024 and since published in the Journal of Empirical Legal Studies. Testing the leading commercial legal AI research tools on hundreds of open-ended queries, the researchers found that Lexis+ AI and Ask Practical Law AI produced incorrect or misgrounded information more than 17% of the time, and Westlaw's AI-Assisted Research more than 34% of the time. The tools marketing themselves as eliminating hallucination did not eliminate it.

The study's most important contribution is the distinction between two kinds of error:

  • Incorrect answers — the model states the law or the fact wrongly.
  • Misgrounded answers — the model states a correct rule but cites a source that does not support it.

Misgrounded answers are the more dangerous failure, and not because they're more frequent. A citation that exists — a real section of a real purchase agreement, a real opinion — disarms skepticism. The reasoning sounds right and the paper trail looks real, so the flag passes a skim and only fails an actual verification. Add the courts' reaction — sanctions for unverified AI citations, of which Mata v. Avianca (S.D.N.Y. 2023) is the landmark — and the conclusion writes itself: verification is not a courtesy, it is the job.

How AI contract review fails — a taxonomy

The failures cluster in predictable places, most of them inherited from the pipeline described in how AI reads a contract:

Misclassification

A clause labeled as the wrong type — an indemnity parsed as a warranty, a change-of-control trigger read as an ordinary consent requirement. Drafting style varies so much across firms that classifiers misfire on both ends: false positives that flag benign language, false negatives that miss the actual risk clause. A wrong classification poisons everything downstream, because the playbook comparison is only as good as the label.

Missed cross-references and defined terms

Contracts are hypertext. A defined term ("Company," "Knowledge," "Confidential Information") is defined once in the front of the agreement and used across dozens of operative clauses; "notwithstanding Section 4(a)" qualifies what looks like a standalone obligation. If the retrieval system pulls the operative clause without its qualifiers and definitions, the model analyzes the term in a vacuum — and confidently reports the wrong answer.

OCR and chunking corruption

Scanned contracts introduce character errors — a dollar figure misread by an order of magnitude, a negative particle dropped. Chunking can split a clause mid-sentence, separating an obligation from its exception. The model reads the corrupted fragment literally and reasons confidently from it.

Misgrounded flags

The flag exists, the classification is right — but the cited passage does not actually support the claim made about it. This is the contract-review version of the Stanford finding, and it is why "open the flag and read the source" is non-negotiable.

Presence over judgment

The machine reliably reports whether a carve-out exists. It cannot reliably tell you whether the carve-out is adequate for this deal — whether the IP indemnity actually covers the infringement risk your client faces, or whether the cap, the indemnity, and the insurance policy combine into a coherent risk position.

Context limits on long documents

Very long agreements with schedules and exhibits push models into what researchers call the "lost in the middle" problem — the critical limitation buried in a long schedule can be down-weighted or missed entirely, even in tools with large context windows.

"The machine's failure is rarely a loud mistake. It is a quiet, plausible one — a wrong label, a missing qualifier, a citation that almost supports the claim."

Why "completeness" is not "understanding"

The entire value proposition of AI contract review is completeness — nothing gets missed, everything gets compared. That is real, and it is the reason the tools earn their place. But completeness is coverage, not comprehension. A complete list of every clause and deviation is not a review memo; it is the raw material for one. The firm that treats the exception list as the finished work product has skipped the step where the law actually happens: reading the flagged clauses in context, against the transaction, against the jurisdiction.

What still needs a human eye

Concretely, the judgment that cannot be delegated:

  • Intent and commercial context. What the parties meant, and whether a "deviation" is actually the point of the deal.
  • Enforceability under governing law. Whether a clause survives the jurisdiction — anti-indemnity statutes, UCC versus common law, local practice — the facts that void a clause in one state and enforce it in another.
  • Cross-clause interplay. How the indemnity, the cap, the insurance, the force majeure clause, and the termination rights work as one risk position — the analysis covered in depth in our look at the clauses AI flags.
  • Negotiation strategy. How hard to push back, what to concede, when to walk away.
  • Definitions, schedules, and exhibits. The operational appendices where the risk actually lives and where models are most likely to lose the thread.
  • Materiality. Which flags matter in this deal — the final call that turns a list into a recommendation.

The verification protocol: how to catch what the AI missed

Verification is not a vague intention; it is a procedure. The version that survives both audit and ethics review:

  1. Open every flag against the source. Confirm the cited passage exists and actually says what the flag claims. This is the direct counter to misgrounding.
  2. Confirm the classification. Is the clause really the type the tool labeled it? One wrong label re-frames the whole comparison.
  3. Read in context. Pull the cross-references and the definitions for every term that qualifies the flagged clause.
  4. Spot-check the unflagged sections. Sample a read-through of sections the tool did not flag — even a small percentage catches false negatives, the errors of omission that a flag list can never reveal.
  5. Inspect scanned documents. For OCR-heavy files, verify the figures and the operative language against the image.
  6. Log the misses. Record every flag that was wrong and every section where the tool stayed silent and missed something. This is the raw material for the playbook review described in building a clause playbook — the machine's misses are the best evidence of where its rules need tightening.

This protocol is what ABA Formal Opinion 512 demands in operational form: the lawyer who uses the tool remains responsible for understanding its limits (Rule 1.1), for supervising everyone who touches it (Rules 5.1 and 5.3), and for protecting the client's confidential information while it's in the tool (Rule 1.6). Delegating the reading does not delegate the responsibility.

The practical economics of the limits

There is a reason the efficiency claims deserve scrutiny. Large firms that piloted enterprise AI reported that auditing the machine's output was so labor-intensive it eroded much of the time savings — several prominent firms (including Paul Weiss) spent over a year testing enterprise products before they could demonstrate hard productivity metrics. The lesson is not that the tools fail; it is that verification is part of the cost of using them, and the firms that budget for it honestly get the real benefit — coverage and consistency — while the firms that skip it get the risk without the savings.

How the honest division of labor works

Put the four posts of this series together and the picture is complete: the pipeline that reads the contract, the clauses that get flagged, the playbook that sets the standard, and now the limits that set the boundary. The machine scans, classifies, and compares; the lawyer verifies, judges, and decides. Neither step works without the other, and the boundary between them is not blurry — it is the checklist above.

How Lawyer Assistant approaches this

Lawyer Assistant is designed around this division of labor. Every finding in its compliance playbook scan links to the exact source text, so Step 1 of the protocol — open the flag against the source — is a click rather than a search. Because everything runs locally on your machine, the Rule 1.6 half of the equation is handled by architecture, not by promise. The tool does the coverage work; the verification protocol and the judgment stay with you.

The bottom line

AI contract review fails in predictable, documented ways — misclassification, missing qualifiers, misgrounded citations, and a quiet confidence in all of them. None of those failures is disqualifying, because all of them are catchable by a disciplined verification step. What would be disqualifying is treating the machine's output as finished work. The tools are excellent readers. They are not yet lawyers — and the gap between the two is exactly the work that still needs a human eye.

"The tool guarantees nothing gets missed. The lawyer guarantees nothing gets believed without checking."

Sources & further reading

  • Magesh, Surani, Dahl, Suzgun, Manning & Ho, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (Journal of Empirical Legal Studies, 2025) — Lexis+ AI and Ask Practical Law AI >17% incorrect or misgrounded; Westlaw AI-Assisted Research >34%.
  • Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023) — sanctions for filing fabricated AI-generated citations.
  • Park v. Kim, 91 F.4th 610 (2d Cir. 2024) — referral to the grievance panel for a hallucinated citation.
  • ABA Formal Opinion 512, "Generative Artificial Intelligence Tools" (July 29, 2024) — competence, confidentiality, candor, and supervision (full text (PDF)).
  • Reported large-firm AI pilot friction (e.g., Paul Weiss) on the labor cost of output verification.

This article is general information about technology and professional practice. It is not legal advice for any specific matter, and rules vary by jurisdiction — verify against the authority applicable to your matter.

Questions, answered

The key questions from this article, answered plainly.

Where does AI contract review still fail?

In a few predictable places: classifying a clause incorrectly, missing cross-references and defined terms, inheriting OCR or chunking errors, and flagging a clause when the cited passage does not actually support the flag. The 2024 Stanford study of leading legal AI tools found incorrect or misgrounded answers in more than 17% of test queries, with one major tool above 34%.

What is the difference between an incorrect and a misgrounded AI answer?

An incorrect answer states the law or the fact wrongly. A misgrounded answer states a correct rule but cites a source that does not support it — the reasoning sounds right, but the paper trail is wrong. Misgrounded answers are the more dangerous failure because they survive a skim and only fail a real verification.

What should a lawyer never delegate to AI in contract review?

The judgment calls: what the parties intended, whether a clause is enforceable under the governing law, whether a deviation is material in this deal, how hard to push back in negotiation, and the final sign-off. Courts have sanctioned lawyers for treating AI output as verified — Mata v. Avianca (S.D.N.Y. 2023) is the landmark example — and ABA Formal Opinion 512 requires verification before reliance.

How do you catch what the AI missed?

Verify every flag against the source passage, confirm the cited text actually supports the claim, check that the clause was classified correctly, and read in context including cross-references and definitions. Then sample the unflagged sections — a read-through of even a small percentage guards against false negatives.

Filed under Practice notes · Contract review · AI reliability ← All articles
Next steps

Put it to work on your own documents.

Lawyer Assistant runs entirely on your machine — install it in minutes, read the documentation, or browse more notes from the Legal Desk.