AI Citation Hallucination Statistics (2026 Data)
First-party outcomes from 47,098 checked references. The original fabrication-rate interpretation was withdrawn after a coverage audit.
Daniel Jyoji Nichiata
Founder & Lead Developer
Check your own references against the same source records
Paste your bibliography and review identifier matches, field differences, and references with no returned match. Try a check without creating an account.
Key findings at a glance
In the 30 days ending July 8, 2026, people submitted 47,098 references to the CiteMe Reference Checker across 2,149 checks. Each reference was compared with records returned by the checker’s active sources. Here are the resulting match outcomes:
- 34.0% of all checked references (15,992 of 47,098) could not be matched to an acceptable source record.
- 90.5% of full bibliographies contained at least one reference that could not be verified at all.
- Only 53.9% of references matched a real, findable work cleanly; 12.2% matched a real work but with wrong metadata (year, authors, or journal); 34.0% could not be matched to any record.
- The median checked bibliography had 17 references.
Where this data comes from
These statistics are first-party data from the CiteMe Reference Checker and Citation Checker. Users paste a reference list (or upload a document), and each entry is checked against scholarly databases — OpenAlex, Crossref, Semantic Scholar, PubMed, Europe PMC, and others — by title, authors, year, journal, DOI, and other identifiers. The numbers above aggregate every check completed in the 30-day window ending July 8, 2026, counted from server-side verification logs rather than optional browser analytics, so consent settings and ad blockers do not skew them.
All figures are aggregates. No individual bibliography, reference, document, or user is identifiable from this data, and cells below a minimum sample threshold are suppressed in our reporting pipeline.
Correction: unverified is not fabricated
On August 10, 2026, a production coverage audit found that the historical fabrication-risk heuristic counted two signals that were not independent: complete citation metadata and a routine six-source search cascade. As a result, a complete reference missed by that cascade could be labeled likely fabricated by construction. The earlier 27.2% reference-level and 84.4% bibliography-level fabrication interpretations are withdrawn.
The verification outcomes remain useful: 34.0% of submitted references had no acceptable match in the connected sources. That group requires human review and can include real books, regional scholarship, legislation, government reports, grey literature, parser failures, and fabricated citations. It cannot support a fabrication-rate estimate on its own.
How accurate is the flag? A labeled benchmark
In July 2026 we built a frozen benchmark of 369 references with ground truth known by construction: 160 correctly-cited real papers (DOI confirmed in CrossRef and OpenAlex), 100 real papers with exactly one deliberately corrupted field (wrong year, misspelled author, wrong journal, altered DOI), 80 constructed nonexistent references (plausible authors, real journals, invented titles — verified absent from CrossRef and OpenAlex at build time), and 29 real but unindexed works — theses from institutional repositories, each with a public repository record as proof of existence. The corpus spans English, Portuguese, Spanish and German, five citation styles, and deliberately sloppy hand-typed variants.
Results of running the full corpus through the checker still support the existence-check claim: not one of the 80 fabricated references came back verified — all 80 failed verification. The previously published recall and false-positive figures for the separate fabrication-risk flag were produced by the superseded heuristic and are withdrawn pending a new representative benchmark.
An offline replay of the frozen captured fields confirmed that the former flag result depended on the removed search-count signal. This replay is diagnostic evidence, not a replacement accuracy benchmark; no recall or false-positive claim for the revised flag will be published until a representative labeled evaluation is complete.
What the outcome data can support
The defensible headline is that 34.0% of the 47,098 references checked in this window did not produce an acceptable source match. This measures verification coverage on a self-selected checker population; it does not measure how many references were fabricated.
Of the 1,553 checks that involved a full bibliography (5 or more references), 90.5% contained at least one reference that could not be verified. The historical count of bibliographies carrying the old risk flag remains in the downloadable dataset for reproducibility, but its fabrication interpretation is withdrawn.
If you want to understand how these fake references are produced and how to recognise one by eye, our companion guide on spotting AI-hallucinated citations covers the mechanics; this page is the measurement.
How this compares with published research
Peer-reviewed studies that asked chatbots to produce citations and then verified them report fabrication rates in the same range or higher. Walters and Wilder (Scientific Reports, 2023) found that 55% of GPT-3.5's citations and 18% of GPT-4's citations were fabricated in a controlled test. A 2023 Cureus study of ChatGPT-generated medical content found nearly half of the references were fabricated and most of the rest contained errors. Research on human-written papers, before generative AI, already put general citation error rates at 25-54% — but those were mostly metadata mistakes pointing at real works, not invented sources.
Our production data describes what reaches a reference-checking workflow after some mix of drafting, editing, and copy-pasting. It establishes a high no-match rate in this self-selected population, but it cannot be compared directly with controlled fabrication rates until the revised risk flag has a representative labeled evaluation.
How to cite these statistics
You are welcome to reference this data in articles, guides, teaching materials, and research with attribution. A ready-made citation:
CiteMe (2026). AI Citation Hallucination Statistics: verification outcomes for 47,098 checked references, 30 days ending July 8, 2026. https://citeme.app/learn/citation-hallucination-statistics
The underlying aggregates are also available as a machine-readable JSON dataset (CC BY 4.0) via the Download Dataset button on the chart below. We plan to refresh these figures periodically as checking volume grows; the published date above matches the data window described on this page, and revisions will carry an updated date and a new dated dataset file.
Verification outcomes for 47,098 checked references
Every reference submitted to the CiteMe Reference Checker in the 30 days ending July 8, 2026, grouped by verification outcome. A later audit invalidated the original interpretation of a historical risk-flag subset.
Ready to cite your sources?
Find source records, inspect their origin, and format citations in the style you need.
Related Articles
7 Best Free Citation Generators for Students (2026)
Compare ZoteroBib, Scribbr, MyBib, CiteMe, EasyBib, BibGuru, Citation Machine for APA, MLA, Chicago, and PubMed/PMID citations. Free tool comparison.
9 Best Free Citation Checkers for Students (2026)
Compare the top free citation checkers: tools that scan your reference list for formatting errors, missing fields, and AI-hallucinated references. Covers CiteMe, ReciteWorks, CiteTrue, CiteSure, Trinka, Scribbr, Grammarly, Citation Format Checker, and PerfectIt.
Citation Generators: How They Work and How to Use Them
Learn how citation generators work, their benefits and limitations, and how to verify automatically generated citations.