AI Flags 250,000 Cancer Studies as Possibly Fake: AI vs AI

A large-scale analysis published in The BMJ reports that a machine-learning “scientific spam filter” flagged nearly one in ten cancer research papers published since 1999 for suspicious writing patterns, underscoring growing concerns over paper mills and automated text generation in scientific publishing.

Study Overview

The study, led by Adrian Barnett, a biostatistician at Queensland University of Technology, deployed a BERT-based classifier trained on 2,202 retracted paper-mill papers. Researchers applied the model to 2.6 million cancer-related studies published between 1999 and 2024, screening for linguistic signals commonly associated with fabricated or mass-produced manuscripts.

Key Findings

  • The tool flagged 261,245 papers—9.87% of the dataset—as exhibiting suspicious writing patterns.
  • The share of flagged papers increased over time, rising from about 1% in the early years of the dataset.

The authors note that a flag indicates atypical or formulaic language features consistent with paper-mill output but does not on its own prove fraud or misconduct. Any flagged paper would require human review and additional evidence to confirm issues.

Why It Matters

The findings highlight scale and trajectory concerns around the integrity of scientific literature, particularly as automated text tools and organized paper mills become more sophisticated. Automated screening systems like the one described could support editors, reviewers, and funders by prioritizing high-risk submissions for closer scrutiny, though expert verification remains essential.

×