Jul 20, 2026
AI

AI detectors falter when models copy a writer’s style

Epoch AI found Pangram, GPTZero and Originality.ai were strong on basic AI text but missed more style-mimicking output, especially scientific writing.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

AI detectors falter when models copy a writer’s style
Photo: The Decoder

Epoch AI tested three widely used AI-writing detectors and found that Pangram, GPTZero and Originality.ai performed well on basic model output but became less reliable when models were told to imitate a specific author. For education and research settings that treat detector scores as evidence, the gap is material: the study found that 13% of style-mimicking AI passages went unflagged on average, with scientific writing producing the highest miss rates.

The research group evaluated Pangram version 3.3.2, GPTZero model 2026-05-11-base and Originality.ai Turbo 3.0.2. Its dataset included 495 human-written passages from 99 authors, split evenly among blogging, fiction and scientific writing. Epoch AI said all of the human texts were written before ChatGPT was released in November 2022, reducing the chance that the sample already contained AI-generated prose.

On straightforward AI generation, the detectors were close to perfect by the study’s measure. Epoch AI reported that the highest false-negative rate, meaning AI text classified as human, was 0.7%. On human writing, Pangram and GPTZero produced no false positives in the test. Originality.ai labeled 19 of the 495 human passages as AI-generated, a 3.8% false-positive rate.

The results changed when the prompts included author samples. Epoch AI used Claude Opus 4.8, GPT-5.5 and Gemini 3.1 Pro, giving each model five real passages from an author and asking it to produce new text in that style. Across 297 generated passages, an average of 38 were not detected. Pangram missed 10% of the style-imitation set, GPTZero missed 11% and Originality.ai missed 18%.

The genre split is the sharper finding. For fiction, the false-negative rate across the detectors ranged from 1% to 5%. In scientific writing, Pangram failed to identify 25% of style-mimicking AI passages, GPTZero missed 24% and Originality.ai missed 29%, according to Epoch AI. The published data showed worse results in some model-and-genre combinations: Pangram missed 48% of Gemini-generated academic passages, while Originality.ai missed 39% of academic passages generated by GPT-5.5.

The three products use different approaches. Pangram relies on a neural network trained on human and machine writing, and its founder has described the system’s decisions as a black box. GPTZero looks at predictability and variation in word choice, based on the premise that model-written text tends to be more uniform. Originality.ai says it detects statistical patterns learned from training data containing human and AI text.

Epoch AI’s findings suggest that strong performance on plain prompts does not translate cleanly to adversarial or style-conditioned writing. They also cut against a narrow reading of detector quality based only on false positives. A tool can avoid accusing human writers in a benchmark and still miss a meaningful share of AI-generated work when the generation process is designed to resemble a real author.

This story draws on original reporting from The Decoder.

More from AI

All AI →