Are AI Detectors Accurate? False Positives and Bias
Vendors publish false positive rates under 1 percent. Independent researchers testing the same detectors on non-native English writing found rates above 60 percent. Here is what the evidence shows.
Turnitin states that its false positive rate stays under 1 percent. Independent researchers who tested a wider set of detectors against essays written by non-native English speakers found an average false positive rate above 60 percent on that same kind of writing. Both figures are published, both are about detection technology sold to the same market, and both are true. So are AI detectors accurate? The honest answer depends on which population was tested, which detector was used, and how the vendor defines the number in the first place.
This post will not re-explain how a detector computes that score. The mechanism, perplexity, token probability, how the classifier is trained, is covered in detail elsewhere, and it is worth reading first if you have not. What matters here is what happens once the score exists: how often it is wrong, whose writing it is wrong about most often, and what a false accusation actually costs someone who did nothing wrong.
Are AI detectors accurate?
Not consistently, and the gap between marketing claims and independent testing is well documented. In 2023, a team led by Debora Weber-Wulff tested 12 detection tools against a mixed corpus of human and ChatGPT-generated text and concluded that the available tools were neither accurate nor reliable, with a built-in bias toward calling machine text human rather than the other way round. The two best performers in that comparison, Turnitin and Compilatio, reached only about 80 percent accuracy, and every tool lost ground once the text had been paraphrased.
Two years on, the picture is not static. A 2025 working paper from the University of Chicago's Booth School, by Brian Jabarian and Alex Imas, tested newer commercial detectors, Pangram, GPTZero and Originality.ai, alongside an open-source baseline, against nearly 2,000 human-written passages spanning blogs, news, reviews and fiction. In that controlled setting, the strongest tool never fell below 99.8 percent accuracy and kept its false positive rate near zero, while the open-source baseline performed close to random guessing. Detectors have genuinely improved. Some of them.
But the catch is the word controlled. Benchmark passages are clean, complete and produced under known conditions. A submitted essay is none of those things reliably. Real world results may not be identical with lab results, as Turnitin's own chief product officer has said publicly, though that's a rare and useful admission from a vendor. The next section sets the vendor numbers next to the independent ones so the gap is visible rather than asserted.
Vendor claims versus independent testing
The table below lines up what four detectors claim for themselves against what independent researchers measured when they tested the tools directly. Read the two middle columns together, not the vendor column alone.
| Detector | Vendor-claimed accuracy | Independently observed | What the vendor says about false positives |
|---|---|---|---|
| Turnitin | 98% confidence threshold before flagging; under 1% document-level false positives | About 80% accuracy, the best of 12 tools in a 2023 comparison, on an earlier version of the tool | Under 1% at document level above a 20% AI-written threshold; about 4% at sentence level |
| GPTZero | 95.7% AI-text detection at a 1% false positive rate on the independent RAID benchmark | 96% accuracy on short passages; false positive rate under 1% in a 2025 University of Chicago Booth study | Reports a 1% false positive rate on the RAID benchmark |
| Originality.ai | 99% accuracy; 0.5% false positive rate (Lite model), 1.5% (Turbo model) | False positive rate under 1%, but missed 10% to 40% of AI-written text in the same Booth study | Publishes separate false-positive figures by model tier |
| Pangram | Error rates it says are over 38 times lower than competing tools; states no bias against non-native English writers | Accuracy never below 99.8%, false positive rate near zero, in the Booth study | States near-zero false positive rates across 10 text domains and 8 language models |
Two things stand out. First, all vendor figures are from a lab that's controlled by the vendor and the Booth figures are from researchers who have nothing to sell. Second, Turnitin does not feature at all in that middle column as the 2025 study did not test it. In the table, its independent figure is based on an earlier 2023 comparison but was done on an older version of the tool, so remember this when looking at the table as a like-for-like comparison.
How accurate is Turnitin's AI detection?
Turnitin does not publish one accuracy figure. It publishes a threshold and a trade-off instead. The company has said it only flags a document when it is 98 percent confident the flagged portion is AI-written, and that this threshold is what keeps the document-level false positive rate under 1 percent. The cost of that caution is recall: Turnitin has estimated it catches roughly 85 percent of AI-assisted writing, letting the rest through deliberately rather than risk a wrong accusation.
At the sentence level, where an interface highlights specific lines rather than a whole document, Turnitin's own published figure is closer to 4 percent. That is the number worth knowing if a professor screenshots a single flagged paragraph rather than a whole-document score, because a highlighted sentence carries a meaningfully higher chance of being wrong than the document score beside it.
Turnitin has also conceded the gap between lab and classroom directly. After early real-world use produced more false positives than the lab figures predicted, particularly on shorter submissions, the company raised its minimum document length for scoring from 150 to 300 words. That is a vendor adjusting its own product because the published number did not hold up outside the lab, which says as much as any independent study.
Humanize your own paper
Transform your AI-assisted text and make it sound human, without touching important words or citations.
What a false positive actually costs a student
An executive MBA student at Yale's School of Management found out in the most literal way possible. A teaching assistant flagged his final exam in a finance course because the writing was, in the TA's own words, long and elaborate with near-perfect punctuation and grammar. A professor ran it through GPTZero, the answers were scored as likely AI-generated, and the student was charged with a violation.
He failed the course and was suspended for a year. His lawsuit, filed in federal court in February 2025, alleges that Yale's own policy restricts the use of tools like GPTZero for exactly this reason, their documented false positive rate against non-native English speakers, and that the tool was used against that policy. As evidence, his filing notes that GPTZero returned a 100 percent AI-generated score on academic writing by Yale's own former university president and the business school's dean.
A judge declined to grant an early injunction in May 2025, and the case remains unresolved as this is written, so treat the allegations as allegations. The plaintiff, a French national, argues that the same documented bias against non-native English speakers explains why his writing scored the way it did, which is the claim the next section examines directly. What is not in dispute is the shape of the cost: a year out of a degree program, a failing grade, legal fees, and months spent contesting a score rather than finishing a degree. A 1 percent false positive rate sounds small until a large population makes it someone specific.
Why do AI detectors misclassify non-native English speakers?
In the study that first put a number on it, Stanford researchers ran seven widely used GPT detectors against 91 TOEFL essays by non-native English speakers and 88 essays by US eighth-graders. The detectors read the eighth-grade essays correctly almost every time. On the TOEFL essays, written by adults fluent enough to sit a graduate-level English exam, the same detectors returned an average false positive rate of 61.3 percent. One detector flagged 97.8 percent of them as AI-generated, and all seven agreed on nearly a fifth of the set.
The mechanism is the same one that governs the rest of detection: predictable vocabulary and simpler sentence construction read as low perplexity, and low perplexity reads as machine-made, regardless of who is holding the pen. Writers working in a second language lean on well-established phrasing because that is what fluency in an additional language actually looks like, and that is precisely the profile a detector was built to flag.
Writers in that position are not well served by advice to simply write more naturally, since natural is exactly what got flagged. TextPulse's page for ESL academic writers covers the phrasing patterns behind that risk in more depth, including what changes them without changing what the sentence means.
The 2023 finding was not a one-off either. A December 2025 benchmark built specifically to test detector bias, covering more than 200,000 samples across demographic and language categories, still found consistently lower recall for English language learners across the open-source detectors it tested. Two years and several model generations later, the bias that Stanford documented had not fully closed.
Not every vendor tells the same story here. Pangram's own technical documentation states that its classifier is not biased against non-native English speakers, a claim the Chicago Booth researchers did not independently test but that at least signals the newer generation of detectors is being built with the problem in mind rather than ignoring it.
If you want to see where your own writing sits before anyone else scores it, a free perplexity checker gives you the same underlying number a detector would compute, without submitting anything to an institution first.
Who else gets flagged, and why
Non-native English speakers carry the largest documented risk, but the same statistical logic catches other writing too. Weber-Wulff's team found that every one of the 12 tools they tested lost accuracy once the text had been paraphrased, which describes a meaningful share of academic writing that has been through a proofreader, an editor, or a supervisor's red pen before anyone runs a detector on it. A paper revised by several co-authors carries a version of the same risk: many people's phrasing averaged toward whatever reads as standard in the field.
None of this means the number is meaningless, only that it needs corroboration before it means anything about a specific person. If you have a text that reads as uniform, similar sentence lengths, similar rhythm throughout, it will score as more machine-like regardless of who wrote it or why. This can be checked directly using a free burstiness checker instead of guessed at.
What a defensible AI detection policy looks like
Yale already had a policy that reportedly restricts tools like GPTZero, for close to the reasons this article has laid out: a documented false positive rate and a documented bias against non-native English writers. The case above is therefore less a story about one unlucky student and more a story about an existing rule not being followed. A score becomes one input rather than a verdict only when a policy actually enforces that distinction.
- Treat the score as one input, never the sole basis for a misconduct finding
- Disclose the tool's published false positive rate to students before using it against them
- Apply extra caution to short submissions and non-native English writing, both documented weak points
- Guarantee a path to appeal that does not require disproving a statistic
None of this is exotic. It is closer to how any single piece of forensic evidence should be treated anywhere else: useful, checkable, and never sufficient by itself.
For an individual writer, the same principle applies before a score exists rather than after. A single number, from any tool, is an estimate rather than a verdict, and treating it as certainty in either direction, safe or caught, asks more of the number than it can support.
TextPulse's own AI humanizer tool reports a Human Score on that same logic: an estimate computed from your own text, not a detector's verdict on it, useful as a signal of how a draft currently reads rather than a promise about what any specific detector will say.
The number on any report will keep moving as vendors retrain their models and researchers keep testing them against harder cases. But what should not move is the habit underneath it: read a score as one measurement among several, ask what population it was validated on, and ask what the vendor is not saying about the cases it still gets wrong.
Frequently Asked Questions
Yes. Weber-Wulff and colleagues found AI-detection tools 'neither accurate nor reliable' in a 2023 comparison, and non-native English writers face the highest documented risk, with one study finding an average false positive rate above 60 percent on TOEFL essays. A single score is not proof of anything by itself, which is why it should never be the only evidence behind an accusation.
PhD in natural language processing, with years spent building NLP applications end to end. Moe works on text analysis: lexical and syntactic structure, and what separates machine-generated prose from human prose statistically. He has been experimenting with computational linguistics since the early days of NLTK, spaCy and WordNet, and still writes most of his tooling in Python.