Open Editor
Privacy
7 min read

Why AI Detectors Keep Flagging Human-Written Essays as AI

A Stanford study found AI detectors falsely flagged 61% of essays by non-native English speakers as AI-generated — every one of them written entirely by hand. The mechanism behind that bias is worth understanding before it costs someone a grade.

FindingDetail
False positive rate, non-native English essays61.2-61.3%
False positive rate, native English essays (same study)Near zero
Essays flagged by at least one of 7 detectors89 of 91 (97.8%)
Essays unanimously flagged by all 7 detectors18 of 91 (19.8%)
False positive rate after adding more elaborate vocabularyDropped to 11.6%

A 2023 Stanford study, still widely cited as the reference point in 2026, tested seven popular AI detectors against 91 TOEFL essays written by non-native English speakers — every single one written entirely by hand. The detectors falsely flagged 61.2% of them as AI-generated on average, while the same tools correctly classified essays from native English speakers with near-perfect accuracy.

Why This Keeps Happening, Not Just Once

Most AI detectors measure two properties of text: perplexity (how predictable each word choice is) and burstiness (how much sentence length and structure vary). Non-native English speakers, especially those educated in formal academic settings, tend to write with more standard vocabulary and more consistent sentence structure than native speakers — exactly the low-perplexity, low-burstiness pattern these detectors are built to flag as machine-generated. The bias isn't a bug in one tool; it's baked into how the underlying detection method works.

The Consequences Have Been Real, Not Theoretical

A student at Liberty University received failing grades on three assignments after an AI detector flagged them, even after showing her revision history and a handwritten first draft of one paper — the grades stood regardless. Separately, a Yale School of Management student and a University of Michigan student have both pursued legal action after AI-detection flags led to disciplinary consequences, citing wrongful accusation and discrimination against non-native English speakers specifically. These aren't isolated incidents; they follow the exact pattern the original Stanford research predicted.

The detectors aren't distinguishing "a machine wrote this" from "a human wrote this" — they're distinguishing predictable writing from unpredictable writing, and treating the first category as suspicious. Formal, simplified, second-language English reads as predictable for reasons that have nothing to do with who wrote it.

A Historical Illustration of the Same Flaw

In July 2023, several AI detectors flagged sections of the United States Constitution — written in 1787 — as AI-generated, simply because its formal, structured legal language happened to match the low-perplexity pattern detectors associate with machine output. The example is old, but the underlying mechanism it demonstrates hasn't changed: formal, technical, or highly structured writing of any origin can trigger the same false flag as second-language English.

Have Newer Detectors Fixed This?

Detection companies report improvement — one vendor's own technical documentation claims a 0% false positive rate on the same TOEFL benchmark used in the original Stanford study, and another reports single-digit rates after updates specifically targeting this bias. These are vendor-reported figures rather than independent, third-party replications, so they're worth treating as a claim to verify rather than a settled fact — the bias that produced the original 61% figure was consistent across seven independently-tested tools, not a flaw unique to one product that a single company's fix would resolve industry-wide.

What Actually Helps If This Happens

Keeping drafts, revision history, and any notes made while writing creates evidence independent of what a detector reports, since detection scores alone have repeatedly been treated as insufficient proof in disputed cases. Running the same text through more than one detector is also worth doing before accepting any single tool's verdict, given how much disagreement exists between them — a text flagged by one tool and cleared by several others is a meaningfully different situation than one every tool agrees on.

The Broader Point Worth Remembering

No independent, peer-reviewed research has shown a text-only AI detector that eliminates this bias entirely, and a detection score has increasingly been treated by courts and institutions as a starting point for investigation rather than proof on its own. Understanding that a flag reflects a statistical pattern in writing style — not a verified fact about who wrote something — is the difference between treating a false positive as a minor inconvenience and treating it as an unchallengeable verdict.


The 61% false positive rate Stanford documented in 2023 remains the reference point because the underlying mechanism — predictable writing reading as machine-generated — hasn't gone away, even as individual tools have improved. For non-native English speakers and anyone whose writing leans formal or structured, that means a detection flag is a signal worth checking, documenting, and disputing, not a fact worth accepting at face value.

For questions or inquiries contact us at info@cleartexteditor.com