Skip to content
Use CaseLegal
Xither Staff13 min read

Government & Professional Services · Use-case guide

Legal AI in Practice: Research, Contracts, Intake, Billing, and the Vendor Landscape

Legal AI works where outputs are checkable and fails where they are citable. Peer-reviewed testing found leading AI legal research tools hallucinate between 17% and 33% of the time[^arxiv-2405-20362], so the deciding factor in every legal AI deployment — research, contracts, intake, billing — is the verification workflow you wrap around the model, not the model itself.

17–33%

Hallucination rate of leading RAG-based AI legal research tools in the first preregistered empirical evaluation — despite vendor claims of "hallucination-free" citations[^arxiv-2405-20362].

Stanford RegLab / HAI, arXiv:2405.20362

58–88%

Frequency of legal hallucinations by general-purpose chatbots asked specific, verifiable questions about random federal court cases — 58% for ChatGPT 4, 88% for Llama 2[^arxiv-2401-01301].

Large Legal Fictions, arXiv:2401.01301

$1,485

Rule 11 monetary sanction a federal court imposed on plaintiffs' counsel in July 2025 for AI-generated citations that used real cases but fabricated quotations and parentheticals[^govinfo-mied-ai-sanctions].

E.D. Mich. sanctions order via govinfo.gov

AI is worth deploying in legal work today — but only in workflows where a human can cheaply verify the output. Contract triage, clause extraction, first-draft documents, and billing review clear that bar. Research memos and briefs clear it only with citation-by-citation checking, because even purpose-built legal AI research tools hallucinate between 17% and 33% of the time[1].

That single empirical fact should anchor how a general counsel, legal operations lead, or CIO structures every legal AI decision. This guide walks the five workflows where legal teams are actually deploying AI — research and drafting, document automation, contract analytics, intake and triage, and billing — and closes with a vendor-evaluation framework and a deployment checklist. Throughout, the question is never "can the model do it?" but "what does it cost to check?"

By the numbers

162

Hand-crafted tasks in LegalBench, the collaboratively built benchmark covering six types of legal reasoning, used to evaluate 20 open-source and commercial LLMs[^arxiv-2308-11462].

LegalBench, arXiv:2308.11462

The tension: fluency is not authority

Legal work is unusual among enterprise AI use cases because its core artifact — the cited authority — is exactly the thing language models are worst at producing. A model that writes a flawless indemnification clause summary can, in the same session, invent a quotation from a real case. The failure mode is not obvious nonsense; it is plausible, confident, correctly formatted error. That asymmetry between how legal AI fails and how buyers expect software to fail is the central tension of the category.

What the pitch impliesWhat the evidence supports
Retrieval grounding "eliminates" hallucination in legal researchRAG reduces hallucination relative to general-purpose chatbots but leading tools still hallucinated 17–33% of the time under preregistered testing[^arxiv-2405-20362]
A real case citation means the quote is realCourts have sanctioned filings where citations were real but the quoted holdings were fabricated — the hardest pattern to catch[^govinfo-mied-ai-sanctions]
Legal reasoning is one capability you can benchmark onceLegalBench needed 162 distinct tasks across six reasoning types to characterize LLM legal performance — capability varies sharply by task[^arxiv-2308-11462]
Lawyer review is a formality in AI-assisted workflowsRule 11 duties attach regardless of good faith; courts have held attorneys must read and confirm the validity of every authority they rely on[^govinfo-mied-ai-sanctions]
The gap between marketing framing and the published evidence base for legal AI.

Why hallucination is the defining fact of legal AI

Two studies frame the problem from both ends. "Large Legal Fictions" tested general-purpose models on specific, verifiable questions about random federal court cases and found hallucinations "alarmingly prevalent, occurring between 58% of the time with ChatGPT 4 and 88% with Llama 2"[2]. The same work found models often fail to correct a user's incorrect legal premises and cannot reliably tell when they are hallucinating[2]. That rules out raw chatbots for any authority-bearing legal task.

The obvious fix — retrieval-augmented generation over a licensed case-law corpus — helps but does not close the gap. The Stanford team behind "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools" ran the first preregistered evaluation of proprietary legal research AI and found that tools from LexisNexis (Lexis+ AI) and Thomson Reuters (Westlaw AI-Assisted Research and Ask Practical Law AI) "each hallucinate between 17% and 33% of the time," even as providers had described their systems as eliminating or avoiding hallucinations[1]. Hallucinations were reduced relative to GPT-4, and the study documented substantial differences between systems in responsiveness and accuracy — but none was safe to leave unchecked[1].

Courts are now pricing this failure mode. In July 2025, a federal judge in the Eastern District of Michigan sanctioned plaintiffs' counsel under Rule 11 after finding briefing citations "with real case citations, but with fabricated quotations or explanatory parentheticals that did not accurately reflect the holding of the case cited," ordering counsel personally to pay $1,485 in defense costs[3]. The court found no bad faith — and imposed sanctions anyway, because Rule 11's duty to read and confirm every cited authority applies regardless of intent[3].

When a case cite is "real," an attorney, or for that matter a judge, might see a case they recognize and assume the quote or holding has been accurately represented[^govinfo-mied-ai-sanctions].
U.S. District Court, E.D. Michigan, sanctions order, July 28, 2025

Note what the court is describing: an evolved, harder failure. As vendors tune models to cite only real cases, the fabrication moves inside the citation — into the quote and the parenthetical — where recognition bias makes reviewers less likely to check. The governance implication is structural: verification cannot sample; it must cover every authority in every filed document. For teams building the control layer around this, the practical techniques live in /guides/hallucination-control-guide, and the liability analysis in /insights/ai-output-risk-and-liability.

Good faith is not a defense

The July 2025 Michigan order held that Rule 11 sanctions "may be imposed regardless of whether an error was made in good or bad faith"[3]. A policy that says "lawyers should review AI output" is not a control. The control is a mandatory, logged, citation-by-citation verification step that a filing cannot bypass.

Workflow 1: Legal research and brief drafting

Research is the highest-value and highest-risk workflow. AI research assistants genuinely compress the retrieval-and-synthesis loop: semantic search across case law and statutes, issue summarization, and first-pass argument framing. But because the output is authority-bearing, the 17–33% hallucination floor documented for even purpose-built tools[1] means the time you save on retrieval is partially spent on verification — and that trade is still usually worth making, if you design for it.

Design the workflow so verification is cheap. First, require the tool to return pinpoint citations with retrievable source text, so a reviewer can check the quote against the opinion in one click rather than re-running the research. Second, separate retrieval from drafting: use AI to find and summarize authority, and treat generated argument text as a template whose every factual and legal assertion gets re-anchored by a human. Third, benchmark candidate tools on your own matters before buying. LegalBench demonstrates why: legal reasoning is not one skill but at least six distinct types across 162 tasks, and model performance varies sharply among them[4]. A vendor's aggregate accuracy number tells you little about the issue-spotting or rule-application tasks your practice actually runs.

Workflow 2: Document automation — NDAs, employment contracts, leases

Document automation is the most mature and lowest-risk legal AI workflow, because it inverts the research problem: instead of generating authority, the system assembles documents from pre-approved language. NDAs, employment agreements, and leases are high-volume, template-driven, and internally standardized — which means the AI's job is selection and adaptation within a clause library your lawyers already blessed, not open-ended generation.

The buying decision here is less about model quality than about template governance. The questions that separate successful deployments: Who owns the clause library, and how are updates versioned and approved? Can business users self-serve standard documents with conditional logic (jurisdiction, deal size, counterparty type) while nonstandard requests route to counsel? Does the tool track which template version produced which executed contract, so a later clause change can be traced across the portfolio? A system that generates beautiful first drafts but cannot answer "which live contracts contain the old limitation-of-liability language?" solves the cheap half of the problem.

Let generation stay boring

For standard documents, prefer deterministic template assembly with AI-assisted intake over free-form LLM drafting. The model's best role in an NDA workflow is classifying the request and extracting the variables — name, term, jurisdiction, scope — not composing confidentiality language from scratch. Reserve generative drafting for the nonstandard documents where a lawyer will review every line anyway.

Workflow 3: Contract analytics — extraction, obligations, risk scoring

Contract analytics runs the pipeline in reverse: instead of producing documents, it reads the portfolio you already have. Three capabilities stack on each other. Clause extraction identifies and isolates provisions — indemnity, confidentiality, termination, assignment — across thousands of executed contracts. Obligation tracking turns extracted commitments (renewal dates, deliverables, payment and reporting duties) into monitored events with owners and alerts. Risk scoring ranks contracts or clauses against your policy baseline so review effort concentrates where deviation is largest.

The reliability profile here is fundamentally friendlier than research. Extraction claims are checkable against the source document sitting next to them — a reviewer confirms a clause classification in seconds, and precision/recall can be measured on your own portfolio before rollout. So pilot on your paper, not the vendor's demo corpus: pull a few hundred representative contracts, hand-label the clauses that matter to you, and score the tool against that ground truth. Accept nothing on faith about accuracy across jurisdictions, legacy formats, and scanned documents, because performance varies with exactly those factors.

Risk scoring deserves the most skepticism. A score is a model opinion about legal exposure, and unlike an extracted clause it has no adjacent ground truth to check. Treat scores as triage signals that order the review queue — never as decisions. Require the vendor to show what features drive a score, keep a human approval step on any action the score triggers, and re-validate the scoring model when your playbook or the regulatory environment changes. An unexplainable risk score embedded in an approval workflow is an audit finding waiting to happen.

Workflow 4: Intake agents and contract triage

Intake is where agentic AI earns its way into legal operations. The pattern: an agent receives inbound documents (email, portal, API), classifies the contract type, extracts counterparty and terms, checks deviation from your standard positions, and routes — standard NDAs proceed on autopilot to signature, everything else lands in a prioritized queue with the deviations pre-flagged. The economics work because NDAs and routine commercial paper are high-volume and low-variance, so even conservative automation absorbs a large share of ticket volume.

Two design rules keep intake agents safe. First, make the autonomy boundary explicit and asymmetric: the agent may approve only documents that match approved templates within tolerance, but may escalate anything — false positives to the human queue are cheap, while a false "standard" that self-executes a nonstandard indemnity is expensive. Second, log every decision with the evidence behind it (which clauses matched, which deviated, what threshold fired), because a triage system whose decisions cannot be reconstructed will not survive its first dispute or regulatory inquiry. Measure the operation with a small set of KPIs — cycle time, straight-through rate, escalation precision — and re-tune quarterly as your templates evolve.

Workflow 5: Billing — time-entry verification and rate analysis

Billing is the quiet workhorse use case: no authority generation, abundant structured data, and a clear counterfactual (what you would have paid without review). AI review of legal invoices splits into time-entry verification — flagging entries that are duplicated, block-billed, vague, or inconsistent with guidelines — and rate analysis, which checks billed rates and timekeeper classifications against negotiated rate cards and historical patterns. Both are anomaly-detection problems where every flag is verifiable against the invoice and the engagement terms, which is why billing review tolerates model imperfection far better than research does.

The adoption risks are relational and procedural, not technical. Auto-rejecting entries on model output alone will poison firm relationships and generate appeals overhead; route flags to a human reviewer with the guideline citation attached, and track the sustain rate of each flag type so you can retire noisy rules. Insist on explainable flags — a rejected line item needs a reason a partner will accept — and an audit trail, because billing disputes are themselves litigable. And normalize your data first: rate analysis across firms only works if invoices arrive in a consistent electronic format with enforced timekeeper classifications.

The ground everything stands on: privilege and confidentiality

Every workflow above sends privileged or confidential material through an AI system, so the data plane is a gating decision, not a feature comparison. Before any pilot, answer four questions in writing. Where is prompt and document data processed and stored, and under what retention terms? Is your data used to train or tune shared models — and is the opt-out contractual, not a settings toggle? Does the deployment model (dedicated tenant, VPC, on-premises) match the sensitivity of the matters involved? And who inside the vendor can see your content, under what access controls and audit logging?

Confidentiality obligations to clients do not relax because a vendor is in the loop, and inadvertent disclosure through a tool's logging or training pipeline is a risk you carry, not the vendor. The general framework for handling personal and sensitive data in AI systems is covered in /guides/personal-data-protection-ai; the discovery-specific handling questions — including how AI-assisted review interacts with privilege logs and production obligations — are treated in the companion piece at /use-cases/ai-ediscovery-litigation-guide.

Map it to a framework you can defend

NIST's Generative AI Profile (NIST AI 600-1, released July 26, 2024) names confabulation as a first-class risk category of generative systems, alongside data privacy and information integrity[5][6]. Anchoring your legal AI controls to its Govern/Map/Measure/Manage structure gives you an externally defensible answer when a client, regulator, or malpractice carrier asks how the tools are supervised.

The vendor landscape, without the logos

The legal AI market is crowded and loud — Harvey, Spellbook, and Ironclad are among the most visible names, the established research platforms from LexisNexis and Thomson Reuters carry AI assistants of their own, and dozens of startups compete in every niche from e-discovery analytics to regulatory intelligence. Resist the urge to shop by name. Vendors in this market resolve into three archetypes, and the archetype — not the brand — determines your integration burden, lock-in profile, and evaluation plan.

ArchetypeWhat you are really buyingPrimary riskEvaluate by
Research-platform AI (assistant bolted to a licensed authority corpus)The corpus and its currency; the AI layer is the interfaceAuthority-bearing hallucination — tested tools of this class hallucinated 17–33% of the time[^arxiv-2405-20362]Blind testing on your own past research questions, scored citation-by-citation
Point solution (drafting, extraction, triage, or billing review)Best-in-class performance on one workflow, integrated into your existing systemsIntegration debt and vendor viability — this segment consolidates rapidlyA ground-truth pilot on your documents; API depth; data-exit terms
Suite / CLM platform (contract lifecycle with embedded AI)A system of record for contracts, with AI features riding on itLock-in — replacing the repository later is a migration project, not a swapThe platform decision on its own merits; treat AI features as roadmap, not guarantees
Three vendor archetypes in legal AI and the evaluation posture each demands.

Across all three archetypes, the same non-negotiables apply: contractual clarity on training-data use and retention; security posture evidenced by current third-party audits rather than logos on a slide; exportable data and models of your annotations so a vendor change does not orphan your work; and reference customers in your practice mix. Pricing in this market is dominated by negotiated enterprise quotes, so model total cost of ownership yourself — licenses, integration, verification labor, and change management — rather than comparing list prices that mostly do not exist.

Honest objections

The strongest case against moving now is that verification overhead eats the gains: if every citation must be checked by hand, a 17–33% error rate[1] means the research assistant is a very expensive way to produce work you must redo. That objection is real for authority-bearing drafting, and teams that cannot fund a genuine verification workflow should confine AI to the checkable workflows — extraction, triage, billing — where it is simply true that review is faster than production.

A second objection: the evidence base is a snapshot, and models improve. Also true — the Stanford study itself notes hallucinations are reduced relative to general-purpose chatbots[1], and vendors iterate quickly. But the direction of improvement is exactly why the Michigan court's warning matters: as systems learn to cite only real cases, residual fabrication migrates into quotations and characterizations, where it is harder to catch, not easier[3]. Improving accuracy changes the error rate; it does not change the governance requirement that every authority be verified. Build the control once and it pays off across every model generation.

Finally, some argue the sanctions risk is overblown — a handful of careless filings, a $1,485 penalty[3]. The dollar figure is beside the point. The exposure is professional (referral to disciplinary bodies), reputational (sanctions orders are public and increasingly reported), and institutional (a client learning its briefs contained fabricated quotes does not parse whose workflow failed). Small sanctions are the warning shots of a norm being set.

The read

Sequence legal AI adoption by verifiability, not by vendor excitement. Start where output is cheap to check and volume is high: intake and triage, document assembly from approved templates, billing review, then contract analytics with a ground-truth pilot. Adopt research and drafting tools last and wrap them in mandatory citation verification from day one, because the published evidence says even the best-in-class tools will hand you fabricated authority a meaningful fraction of the time[1]. Choose vendors by archetype, contract hard on data handling, and anchor the whole program to a defensible framework. The firms that win with legal AI will not be the ones with the best model — they will be the ones whose verification workflow lets them use imperfect models safely at scale.

How to apply this: the legal AI deployment checklist

  • Rank candidate workflows by verification cost: intake/triage, template-based document assembly, and billing review first; contract analytics next; research and brief drafting last.
  • Write the data-plane requirements before any pilot: processing location, retention, contractual training-data opt-out, deployment model, and vendor-side access controls.
  • Pilot every tool on your own materials with hand-labeled ground truth — your contracts for extraction, your past research questions for research assistants — and score results before contracting.
  • Make citation verification mandatory and logged for any AI-assisted document that will be filed or sent externally; check the quote and holding, not just that the case exists.
  • Set asymmetric autonomy boundaries for intake agents: auto-approve only within-template matches, escalate everything else, and log the evidence behind every routing decision.
  • Treat risk scores as queue-ordering signals with a human approval step — never as automated decisions — and require score explainability from the vendor.
  • Classify each vendor by archetype (research platform, point solution, suite) and apply the matching evaluation posture and lock-in analysis before comparing features.
  • Negotiate exportable data and annotations, current third-party security audits, and clear training-data terms into every contract.
  • Map controls to the NIST AI RMF and its Generative AI Profile so supervision of confabulation risk is documented and externally defensible[^nist-ai-600-1].
  • Define KPIs per workflow (cycle time, straight-through rate, flag sustain rate, verification findings) and review them quarterly, re-tuning as templates, models, and case law shift.

Sources

Every quantitative or attributed claim above is linked to a primary source. Last verified at publication.

  1. [1]
    Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
    arXiv (Stanford University authors) · · accessed
  2. [2]
  3. [3]
  4. [4]
  5. [5]
    AI Risk Management Framework
    NIST · accessed
  6. [6]