Independent Research & Regulatory Compliance Monitor
Statutory Benchmarks: NAIC AI Model Bulletin · Cal. Ins. Code § 14021 · Colorado SB 21-169

Technical Autopsy · Evidentiary Integrity

The Anatomy of an AI Hallucination in Bodily Injury Demands

How large language models fabricate non-existent disc herniations, invert radiological negatives, and duplicate medical specials—and how casualty defense adjusters can systematically dismantle synthetic claims in litigation.

The Evidentiary Core: Fed. R. Evid. 901 & State Business Records Exceptions

An AI-generated medical chronology or settlement demand has zero independent evidentiary value in a court of law. Only certified, original medical provider records and itemized billing statements authenticated by record custodians under Rule 803(6) are admissible. When LLM summaries diverge from primary charts, reliance on the summary creates fatal liability for both plaintiffs and insurers.

The Probabilistic Reality Behind Claims AI

In marketing collateral delivered to law firms and TPAs, vendors describe their AI models as “reasoning engines” and “expert medical reviewers.” In reality, all modern commercial AI tools operate on autoregressive transformer architectures: probabilistic mathematical systems designed to predict the next most statistically plausible sequence of words (tokens).

In conversational contexts, small inaccuracies are harmless. But in bodily injury claims adjusting—where the difference between an acute traumatic spinal nerve impingement and age-related degenerative facet arthropathy can alter claim reserves by $250,000—probabilistic token prediction introduces devastating factual distortions.

The Four Most Common Medical AI Hallucination Modes

Through line-by-line audits of AI-generated personal injury demand packets and carrier medical chronologies, we have identified four repeatable failure modes inherent to ungrounded LLMs:

Hallucination Mode Actual Primary Medical Record Fabricated AI Demand Assertion Underlying Model Vulnerability
Negation Inversion “Lumbar MRI: No evidence of acute traumatic fracture or spinal canal stenosis.” “MRI confirms severe lumbar spinal canal stenosis and acute traumatic disc herniation.” LLMs frequently drop linguistic negation modifiers (“no evidence of”) when attention heads focus on high-weight clinical nouns (“stenosis”).
Chronology Conflation “History: Patient reports prior motor vehicle accident in 2018 with ongoing neck pain.” “Client sustained severe, debilitating cervical trauma directly caused by the subject collision.” Failure of temporal attention mechanisms across multi-hundred page PDF medical packets, attributing pre-existing conditions to acute events.
Phantom Diagnostic Multiplier “Assessment: Rule out mild cervical strain. Conservative home exercise recommended.” “Diagnosed with severe chronic cervical radiculopathy requiring future surgical discectomy.” Model treats speculative “rule out” differential diagnoses as verified permanent impairments to match typical demand templates.
Synthetic Billing Duplication Single $18,400 spinal epidural procedure billed on facility master ledger, CMS-1500, and collection letter. Demands $36,800 or $55,200 in total special damages by summing every instance where the number appears in the packet. Lack of cross-document entity deduplication across heterogeneous billing formats.

The Defense Playbook: Dismantling Synthetic AI Demands

When insurance casualty adjusters and defense counsel receive AI-drafted demand packets from plaintiff firms using tools like EvenUp, they should deploy the Four-Point Verification Audit:

  1. Demand Exact Bates Citations: Reject narrative summaries that do not include pinpoint Bates stamps for every asserted injury. If the demand claims a disc herniation, require the exact page number of the radiologist's signed report.
  2. Cross-Audit CMS-1500 Forms Against Ledger Totals: Run strict programmatic deduplication between itemized facility charges, professional physician fees, and collection balances. In over 40% of audited AI demands, special damages are inflated by double-counted line items.
  3. Depose the Plaintiff on AI Assertions: In plaintiff depositions, read specific hallucinated paragraphs from the demand letter and ask the plaintiff whether their treating doctor ever communicated that diagnosis. Invariably, plaintiffs admit they never heard of the severe conditions fabricated by the AI.
  4. Preserve Demands for Bad-Faith Defense: When an insurer declines an inflated policy-limits demand that was substantiated only by synthetic AI summaries, the insurer's line-by-line audit serves as conclusive evidence of a “genuine dispute,” immunizing the carrier from subsequent bad-faith failure-to-settle claims.

The Only Defensible Standard: Verifiable Provenance

The industry's path forward is not to abandon automated indexing, but to enforce strict technological boundaries: systems must never summarize without bidirectional Bates-stamped document provenance. Every claim, dollar figure, and medical finding displayed to an adjuster must link directly and visually to the primary physician signature and original billing receipt.

Research & Editorial Methodology

This report was authored by the Claims Governance Institute Research Group. CGI is an independent, non-partisan research monitor examining artificial intelligence, algorithmic accountability, and regulatory compliance across the insurance and legal-tech sectors. Learn more at our Editorial Standards & Disclosures.