Large language models are being handed the keys to some of the most legally sensitive documents we produce. Before we let them in, we need to understand — and honestly reckon with — how often they lie.
Not long ago, I watched a demonstration of a generative AI system producing a polished quarterly earnings narrative — complete
with revenue figures, margin commentary, and a confident citation to an SEC disclosure. The room was impressed. Then someone quietly crosschecked the numbers. Two figures were wrong. The cited disclosure didn’t say what the system claimed. The regulatory threshold it
referenced had been updated two years earlier. Nobody in the room had caught it at first glance, because everything looked so authoritative.
That experience crystallized for me a problem I’ve been studying closely: hallucination in large language models (LLMs) when they are applied to financial reporting tasks. LLMs are increasingly being used to assist with earnings commentary, risk disclosures, audit narratives, and regulatory filings. The efficiency gains are real. But so is the risk — and I don’t think we’ve been sufficiently honest about its scale.
What We Mean by Hallucination — and Why Finance Is Different
In the natural language processing literature, hallucination refers broadly to content that sounds plausible but is factually incorrect or unsupported by any source.1 In most domains, this is an irritant. In financial reporting, it’s a liability — literally. An incorrect earnings-per-share figure in a public disclosure, a misquoted capital adequacy threshold, or a fabricated reference to an audit standard can expose an organization to SEC enforcement, shareholder litigation, and reputational harm that takes years to repair.
What I’ve observed across my own research — and what colleagues in the space are reporting independently — is that financial reporting tasks are particularly prone to hallucination, and for a predictable reason. The language of finance is simultaneously highly structured and highly technical. LLMs are exceptional at producing fluent financial prose. They know the genre. They know the vocabulary. But fluency is not factual grounding, and the very confidence that makes these outputs useful in drafting also makes errors harder to spot during review.
Four Ways Models Get It Wrong
Through systematic evaluation of LLM outputs across financial report generation tasks — drawing on real annual filings, earnings transcripts, and regulatory disclosure templates — I’ve found that hallucinations in this domain tend to cluster into four distinct patterns, each with its own risk profile.
The Numbers Aren’t Reassuring
When I evaluated a range of frontier LLMs on standardized financial report generation prompts — using real S&P 500 10-K filings as
grounding documents — even the best-performing system produced outputs with identifiable hallucinations in roughly one in seven
responses under zero-shot conditions.2 Less capable models failed at nearly one in three. The gap between the top and bottom performers is real and worth paying attention to, but the more important lesson is that no current model is hallucination-free in this domain — and a 14% error rate in financial disclosure is not a rounding error.
Notably, the sections that performed worst weren’t the narrative summaries — they were the footnote disclosures. Technical footnotes
require precise cross-referencing between accounting standards, company-specific policies, and prior-period comparatives. They are
exactly the kind of content where LLMs produce confident errors that read like authoritative text. This is where I’d urge practitioners to be most cautious.
Mitigation Helps — But Isn’t a Silver Bullet
The good news is that thoughtful deployment architecture can substantially reduce hallucination rates. In my evaluations, retrieval augmented generation — grounding the model’s outputs in retrieved document chunks from the actual
rates by roughly 40%. Prompting strategies that encourage explicit selfverification before output helped further, particularly for numerical
claims. The most effective approach I tested combined structured knowledge retrieval with a lightweight symbolic consistency check on
key financial facts, which pushed error rates below 5% on the same benchmark. That’s not zero, but it’s a different conversation.
The caveat is that each of these approaches adds engineering complexity, latency, and maintenance overhead. Retrieval requires well-maintained document stores. Self-verification requires careful prompt engineering and increases inference cost. Symbolic grounding requires structured knowledge extraction pipelines. None of this is prohibitive, but it means that deploying an LLM for financial reporting is not a drop-in integration — it’s a system design problem.
What This Means for the Computing Community
I want to close with a broader observation, because this isn’t just a financial services problem. The pattern I’ve described — fluent,
confident outputs that are structurally correct but factually unreliable — is a general property of current LLM architectures, and it will surface
wherever we deploy these systems in domains that require precise factual recall. Legal documents, medical records, scientific reports,
engineering specifications. The finance case is instructive precisely because the stakes are visible and the errors are measurable.
As computing professionals, we have a responsibility to build systems that are honest about their limitations — not just in model cards and
research papers, but in how we architect, deploy, and communicate AI systems to the organizations that use them. Hallucination is not a bug that will be patched away in the next model version. It reflects something structural about how language models work.3 The path forward is not to pretend otherwise, but to design for it: with grounding, verification, human oversight, and a clear-eyed understanding of what these systems can and cannot reliably do.
The room I described at the beginning eventually caught the errors — but only because someone checked. That someone matters. The lesson I’ve taken from this work is that in high-stakes domains, “someone checks” needs to be an architectural guarantee, not an afterthought.
REFERENCES & FURTHER READING
1. Ji, Z. et al., “Survey of Hallucination in Natural Language Generation,” ACM Computing Surveys, 2023. doi.acm.org
2. Maynez, J. et al., “On Faithfulness and Factuality in Abstractive Summarization,” ACL Proceedings, 2020.
3. Bender, E. et al., “On the Dangers of Stochastic Parrots,” FAccT 2021. A foundational read on what language models fundamentally are and are not.
4. Lewis, P. et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” NeurIPS, 2020. The grounding architecture
referenced in this article.
5. SEC Staff Bulletin on AI in Investor Disclosures, U.S. Securities and Exchange Commission, 2024.
6. Wei, J. et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,” NeurIPS, 2022.
