Universities Are Quietly Ditching AI Plagiarism Detectors as False Positives Pile Up

EthicsAI in Education
Illustration generated by AI: Editorial image for Universities Are Quietly Ditching AI Plagiarism Detectors as False Positives Pile Up

The Core · TL;DR

  • Reddit testers found AI detectors flagged the 1776 Declaration of Independence as up to 100% AI-generated, underscoring the tools' unreliability
  • Stanford found false positive rates as high as 61.3% on pre-2020 papers; a separate study put GPTZero's false positive rate near 16%
  • Vanderbilt, Dundee and Greenwich have dropped Turnitin, and the UK's higher education ombudsman has warned against using such tools
  • Turnitin says it has cut its false positive rate below 1% by defaulting to 'human-written' when confidence is low, but disputed cases remain unresolved

A test run by Reddit users found that several AI-detection tools flagged the 1776 US Declaration of Independence as 95 to 100 percent AI-generated. The document predates ChatGPT by roughly two and a half centuries, and the result has become a shorthand example for a problem now troubling universities on both sides of the Atlantic: the software schools rely on to catch AI-written coursework is frequently wrong.

Mike Perkins, a researcher at the British University in Vietnam, argues that the false positive rates produced by these tools are too high to justify using them for decisions that can end a student's academic career or damage a researcher's reputation. His concern is not abstract. One case involved an autistic student who received a zero on a research paper after a detector flagged it as AI-written, with the student maintaining that the tool simply misread their natural writing style as machine-generated.

The scale of the underlying data explains why the stakes are so high. Turnitin, launched in its current AI-detection form in April 2023 and used by most universities in the UK, has scanned more than 200 million papers since then. Roughly 22 million of those, about 11 percent, were flagged as containing AI-generated passages, some at levels reaching 20 percent of the text. Those numbers alone justify institutional scrutiny, but the accuracy problem cuts the other way too. A Stanford University test in 2023 ran 91 scientific articles published before 2020, well before generative AI existed, through detection tools and found over half were misclassified as AI-written, with false positive rates as high as 61.3 percent. A separate study found GPTZero caught every AI-generated paper in its sample but also wrongly flagged genuine human writing about 16 percent of the time.

Institutions Are Pulling Back

The response has been uneven but increasingly skeptical. Vanderbilt University stopped using Turnitin after months of internal testing and consultations with both Turnitin staff and outside AI specialists. The universities of Dundee and Greenwich dropped the tool over what they described as insufficient transparency in how it reaches its conclusions. In England and Wales, the higher education ombudsman has gone further, warning universities against relying on these detectors at all after fielding a wave of student complaints.

Turnitin's own leadership has pushed back on how the tool is being used rather than what it reports. N. Chishiteli, the company's head of production, has said the software was never meant to serve as definitive proof of misconduct, but rather to flag material for educators to discuss directly with students. The company says it has since recalibrated its settings, pushing the false positive rate below 1 percent by defaulting to a human-written classification whenever AI authorship cannot be confirmed with high confidence.

That adjustment may reduce the error rate going forward, but it does little to resolve the disputes already generated by earlier, less conservative versions of the software, or to answer the broader question Perkins and others are raising: whether any detection tool available today is reliable enough to underpin punitive academic decisions.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram

More from Research

View all in Research