Aug 8, 2026
AI

AI stories vs human stories: study finds labels and origin both sway ratings

Readers rated ChatGPT short stories above human work overall, but a human byline boosted scores regardless of who wrote the text.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

AI stories vs human stories: study finds labels and origin both sway ratings
Photo: The Decoder

A Villanova University study of AI stories vs human stories found that readers rated ChatGPT-generated fiction more highly than human-written fiction overall, while giving an additional boost to stories presented as human-authored. The result separates the text’s actual origin from the effect of the byline: in this experiment, a human-authorship label was associated with higher ratings regardless of who wrote the story.

The research, by Sydney Sears and Deena Skolnick Weisberg and published online August 5 in Judgment and Decision Making, tested whether readers could judge short fiction’s quality and identify its creator. It does not establish that AI fiction is objectively better, or that the result applies to novels, publishing markets, or all models.

Can readers tell AI stories from human stories?

Generally, not in this study. In the first experiment, 1,682 adults each read one of six stories: three works by published human authors and three corresponding stories generated with ChatGPT. Participants received an attribution saying either a human or AI wrote the story, but the label was sometimes false, then rated its quality and how absorbing they found it.

AI-generated stories received higher ratings for quality and absorption than the human-authored stories. Their strongest scores came when participants believed a human had written them. At the same time, readers rated stories more favorably when told they were written by a person, including when that attribution was inaccurate.

Two follow-up experiments tested recognition without a byline. A total of 424 people in one experiment and 481 in another each read one human story and one AI story on the same theme, then selected which was which. Accuracy was 39.93% in the first and 51.97% in the second, outcomes the researchers said provided no general evidence that participants could reliably distinguish the origins.

Self-reported familiarity with AI systems was associated with better identification performance, whereas self-reported experience with fiction was not. The study examined a small set of short stories generated by one system, ChatGPT, rather than a representative cross-section of literary writing or generative models.

What do the results measure?

The findings measure reader response to particular texts and labels, not a settled measure of literary merit or creativity. The authors proposed that AI prose may have benefited from being clearer, more direct and easier to process, while human fiction may be more subtle or complex. That is their interpretation of the pattern, not a mechanism demonstrated by the experiment.

A separate ACM study points to why the distinction matters. It used expert-developed tests of fluency, flexibility, originality and elaboration on 48 stories, and found LLM-generated stories passed three to 10 times fewer creativity tests than professionally written ones. Its different models, materials and scoring system mean it does not refute the Villanova result. It shows that a favorable reader rating and expert-assessed creativity are different measurements, a distinction relevant to evaluating AI models for the work they will actually do.

For publishers and product teams building writing tools, the narrower takeaway is that disclosure can shape reception even where audiences struggle to identify machine-generated prose. The study does not answer whether readers would make the same judgments in longer works, in editorial settings, or when asked to assess originality rather than immediate quality and engagement.

This story draws on original reporting from The Decoder.

More from AI

All AI →