Efforts to quantify the spread of fabricated references in biomedical journals may be undermined by the very tools used to detect them, according to a new correspondence published in The Lancet. The letter responds to a study by Maxim Topaz and colleagues that reported accelerating rates of fabricated citations in peer-reviewed literature, a trend the authors attribute largely to researchers using large language models (LLMs) when preparing manuscripts. The correspondence argues that the reliability of those estimates hinges on an LLM classifier deployed at a decisive stage of the study's analytical pipeline.

The original research by Topaz and colleagues documented a sharp rise in references that appear plausible but do not correspond to real publications. Such fabricated citations can slip through peer review because they mimic the formatting, author names, and journal titles of legitimate sources. The phenomenon has raised alarms across biomedical publishing, where citation accuracy underpins the credibility of systematic reviews, clinical guidelines, and drug safety assessments. If fabricated references enter the evidence base, they can distort meta-analyses and mislead clinicians who rely on published literature to make treatment decisions.

The correspondence does not dispute that fabricated citations are increasing. Instead, it focuses on the methodology used to measure that increase. The letter notes that the study's estimates depend on an LLM classifier to distinguish genuine references from fabricated ones. If that classifier produces false positives or false negatives, the resulting rates could be overestimated or underestimated. The concern is not merely technical. Detection tools are increasingly being proposed as a first line of defense against citation fraud, and their accuracy determines whether journals, funders, and research integrity offices can trust the numbers they use to justify new policies.

The letter points to a broader challenge facing scientific publishing as generative AI tools become commonplace in manuscript preparation. LLMs can generate references that look authentic, complete with digital object identifiers and journal names, even when no such article exists. This capability has made traditional manual checks insufficient at scale. Publishers have begun deploying automated screening systems, but those systems rely on classifiers and databases that may themselves contain errors or gaps. The correspondence suggests that without independent validation of these detection methods, the scientific community risks building its response to fabricated citations on uncertain foundations.

The implications extend beyond biomedical journals. Research integrity offices, academic institutions, and funding agencies are developing policies to address AI-assisted misconduct. Accurate measurement of the problem is essential for allocating resources, setting editorial standards, and evaluating whether interventions actually reduce the prevalence of fabricated references. If detection tools are unreliable, policymakers may misjudge the scale of the threat or implement measures that fail to address its root causes. The correspondence calls for greater transparency in how LLM classifiers are trained, tested, and validated before their outputs are used to shape publishing policy.

The letter also highlights the difficulty of distinguishing between intentional fabrication and inadvertent error. Researchers under pressure to publish may use LLMs to draft reference lists without verifying each entry, leading to citations that are fabricated in effect even if not by design. This blurring of intent complicates efforts to assign responsibility and design effective prevention strategies. The correspondence argues that prevention must include better training for researchers, clearer journal guidelines on AI use, and routine verification of references at multiple stages of the publication process.

Topaz and colleagues' study has drawn attention to a problem that many editors suspected but few had quantified. The correspondence adds a cautionary note: the urgency of addressing fabricated citations should not lead to premature reliance on detection methods that have not been rigorously evaluated. As LLMs become more capable, the arms race between fabrication and detection is likely to intensify. The letter suggests that the scientific community needs a coordinated approach that combines technological tools with human oversight, transparent reporting of detection accuracy, and shared standards for what constitutes acceptable use of AI in research writing.

The exchange in The Lancet reflects a growing debate about how to safeguard the integrity of the scientific record in an era of rapid AI adoption. Fabricated citations are only one symptom of a wider challenge: ensuring that the speed and convenience of new tools do not erode the trust that underpins scientific progress. The correspondence concludes that measuring the problem accurately is a prerequisite for solving it, and that the classifiers used to measure it deserve the same scrutiny as the manuscripts they are designed to screen.

7Views

Logan Weston

Author

Sports Writer

Logan Weston covers public affairs, politics, business, culture and daily news for Science Official. The role focuses on verification, context, and clear explanations for readers.