AI Detectors Wrongly Flagged Non-Native English Essays
Research finds common detection tools misidentify straightforward writing as machine-generated.
Updated on Sept. 28, 2026 in Artificial Intelligence

Live Poll
Should schools use AI detection scores as the sole evidence to punish students for cheating?
Stanford researchers found that seven AI detectors misidentified over 61% of essays written by non-native English speakers as AI-generated. The findings highlight a significant accuracy gap, as nearly 98% of the student essays studied were flagged by at least one tool.
Why it matters
The systematic misclassification threatens academic integrity protocols by creating a false bias against non-native speakers who employ predictable, simple syntax. Because these detectors rely on statistical patterns, they struggle to distinguish between AI-generated text and straightforward human prose.
In a study of 91 essays, seven detectors flagged 89 as AI-generated; however, rewriting those same texts with more sophisticated vocabulary reduced the average false detection rate to approximately 12%.
The players
Stanford University
An academic institution known for large-scale research into the social and technical implications of artificial intelligence.
OpenAI
A developer of large language models that maintains significant influence over the standardization and oversight of AI generation tools.
Turnitin
A provider of academic plagiarism detection software that provides guidance to educators on the limitations of AI scoring.
Jamir Nazir
A 62-year-old writer based in Trinidad who was falsely accused of using AI tools despite possessing human-written drafts.
The details
AI detectors work by calculating the perplexity—a measure of how predictable a sequence of words is—and the burstiness—the variation in sentence structure—of a text. Non-native English speakers often utilize familiar, straightforward vocabulary and sentence patterns, which the software interprets as the low-complexity output typical of large language models. This creates a statistical overlap where clear, accessible human writing is indistinguishable from machine-generated content to these tools.
Timeline
January 2023: OpenAI launched an AI text classifier.
July 2023: OpenAI discontinued its AI text classifier.
May 2026: Writer Jamir Nazir won a regional Commonwealth Short Story Prize.
The Tech Race
This research follows a pattern set by the discontinuation of OpenAI's AI text classifier due to its failure to meet acceptable accuracy benchmarks. The field continues to struggle with tools that prioritize rapid inference over the linguistic nuance required for reliable attribution.
Educators are now cautioned against relying solely on AI detection scores for disciplinary action. Students and writers concerned about false positives may find that varying sentence structure and vocabulary can reduce the likelihood of erroneous flagging by existing detector software.
The takeaway
The study suggests that current detection reliance on simple statistical predictability will continue to penalize straightforward, clear communication. Readers should monitor future updates from Turnitin and other academic software providers regarding the refinement of their detection metrics to account for English-proficiency variations.
Further reading
For broader context on how these tools are evolving, visit the Artificial Intelligence section.
Live Poll
Should schools use AI detection scores as the sole evidence to punish students for cheating?









