Popular AI text detectors catch plain AI-generated text with near-perfect accuracy. But when language models deliberately copy a specific author's writing style, up to one in five AI texts slips through undetected. Scientific writing is where the detectors fail the hardest.
A research team from Epoch AI tested three of the most widely used AI text detectors: Pangram (version 3.3.2), GPTZero (model 2026-05-11-base), and Originality.ai (Turbo 3.0.2). The test covered three categories: genuine human writing, AI text generated from simple prompts, and AI text that deliberately mimicked a specific author's style.
The team built a corpus of 495 human passages from 99 authors, evenly split across blogging, fiction, and scientific writing. All texts were written before ChatGPT's release in November 2022, which effectively rules out contamination by language models.
When dealing with plain AI-generated text, all three detectors performed almost flawlessly, with the false-negative rate topping out at 0.7 percent. Human texts were also classified correctly for the most part. Pangram and GPTZero didn't produce a single false alarm. Originality.ai, however, flagged 19 out of 495 human passages as AI-generated, a troublingly high false-positive rate of 3.8 percent.











