AI Detection

Are AI Detectors Accurate? False Positives and Bias

Vendors publish false positive rates under 1 percent. Independent researchers testing the same detectors on non-native English writing found rates above 60 percent. Here is what the evidence shows.

Updated on 13 min read
Are AI detectors accurate? Comparison table of vendor-claimed versus independently observed false positive rates for Turnitin, GPTZero, Originality.ai and Pangram.

Turnitin claims that their false positive rate is less than 1%. A group of independent researchers tested multiple detectors on a broader range of writing samples from non-native English speakers. They found the average false positive rate was over 60% on that same type of writing. Both numbers are real, both were published, and both relate to detectors that sell to the same market. Are AI detectors accurate? It depends on what population you tested, what detector you used, and how your vendor defined it in the first place.

This post will not re-explain how a detector computes that score. The mechanism, perplexity, token probability, how the classifier is trained, is covered in detail elsewhere, and it is worth reading first if you have not. What matters here is what happens once the score exists: how often it is wrong, whose writing it is wrong about most often, and what a false accusation actually costs someone who did nothing wrong.

Are AI detectors accurate?

Not consistently, and the gap between marketing claims and independent testing is well documented. A group of researchers led by Debora Weber-Wulff published their findings on 14 detection tools (including 12 free and two commercial) that were tested against a mixture of human and ChatGPT text in 2023. They found that none of these tools are accurate or reliable, with a built-in bias to label machine text as human rather than the other way round. Two of the commercial tools included in this test were Turnitin and PlagiarismCheck. All of the tools in this test performed worse when the text was paraphrased.

But two years on, things aren't stagnant. A working paper from the University of Chicago's Booth School, by Brian Jabarian and Alex Imas, put three newcomers, Pangram, GPTZero and Originality.ai, and one open-source competitor through their paces with almost 2,000 human-written pieces from blogs, news, reviews and fiction. Under controlled test conditions, the best tool in the mix never dropped below 99.8 percent accuracy and held its false positive rate at a low level. The open-source model did as well as a random guesser. Things really have gotten better. Sort of. A little bit.

But the catch is the word controlled. Benchmark passages are clean, complete and produced under known conditions. A submitted essay is none of those things reliably. Real world results may not be identical with lab results, as Turnitin's own chief product officer has said publicly, though that's a rare and useful admission from a vendor. The next section sets the vendor numbers next to the independent ones so the gap is visible rather than asserted.

Vendor claims versus independent testing

The table below lines up what four detectors claim for themselves against what independent researchers measured when they tested the tools directly. Read the two middle columns together, not the vendor column alone.

DetectorVendor-claimed accuracyIndependently observedWhat the vendor says about false positives
Turnitin98% confidence threshold before flagging; under 1% document-level false positivesAbout 80% accuracy, the best of 12 tools in a 2023 comparison, on an earlier version of the toolUnder 1% at document level above a 20% AI-written threshold; about 4% at sentence level
GPTZero95.7% AI-text detection at a 1% false positive rate on the independent RAID benchmark96% accuracy on short passages; false positive rate under 1% in a 2025 University of Chicago Booth studyReports a 1% false positive rate on the RAID benchmark
Originality.ai99% accuracy; 0.5% false positive rate (Lite model), 1.5% (Turbo model)False positive rate under 1%, but missed 10% to 40% of AI-written text in the same Booth studyPublishes separate false-positive figures by model tier
PangramError rates it says are over 38 times lower than competing tools; states no bias against non-native English writersAccuracy never below 99.8%, false positive rate near zero, in the Booth studyStates near-zero false positive rates across 10 text domains and 8 language models

Two things stand out. First, all vendor figures are from a lab that's controlled by the vendor and the Booth figures are from researchers who have nothing to sell. Second, Turnitin does not feature at all in that middle column as the 2025 study did not test it. In the table, its independent figure is based on an earlier 2023 comparison but was done on an older version of the tool, so remember this when looking at the table as a like-for-like comparison.

How accurate is Turnitin's AI detection?

Turnitin does not publish one accuracy figure. It publishes a threshold and a trade-off instead. The company has said it only flags a document when it is 98 percent confident the flagged part is AI-written, and that this threshold is what keeps the document-level false positive rate under 1 percent. Turnitin has estimated it catches roughly 85 percent of AI-helped writing, letting the rest through deliberately rather than risk a wrong accusation.

The actual percentage of false positives that Turnitin returns at the sentence level (where an interface highlights specific lines rather than a whole document) is actually closer to 4 percent. This is the number worth knowing if a professor screenshots a single flagged paragraph rather than a whole-document score, because a highlighted sentence carries a meaningfully higher chance of being wrong than the document score beside it.

It's also worth noting that Turnitin has acknowledged this gap directly. They found that their first real-world results (particularly for shorter documents) had higher rates of false positives than was expected based on their lab tests. So they adjusted their minimum document length to be scored from 150 to 300 words. That's a vendor adjusting its own product because the published number did not hold up outside the lab, which says as much as any independent study.

What a false positive actually costs a student

This happened to an executive MBA student at Yale's School of Management. The student was charged with a violation, in the most literal sense possible. The detector involved was GPTZero. The student's final exam in a finance course had been flagged by a teaching assistant who thought the writing was long and elaborate with near-perfect punctuation and grammar. And this particular professor was running it through GPTZero, which scored the answers as likely AI-generated.

He failed the course and was suspended for a year. His lawsuit, filed in federal court in February 2025, claims that Yale's own policy restricts the use of tools such as GPTZero for just this reason, its documented false positive rate against non-native English speakers, and that he used it against that policy. His filing also cites that GPTZero returned a 100 percent AI-generated score on academic writing by Yale's own former university president and the business school's dean.

The judge refused to issue an early injunction in May 2025. The case has yet to be resolved (as I write). The plaintiff, a French national, claims that the same known bias against non-native English speakers caused his writing to perform the way it did. That's what we look at in the next section. What's not up for debate is the price: a year lost from a degree program, a failing grade, legal expenses, and several months fighting over a score instead of earning a degree. A 1 percent false positive rate sounds little until a big population makes it someone specific.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Which Writing Traits Make Human Text Look Machine-Made

Four kinds of ordinary writing carry the most risk, and none of them involve a person doing anything wrong. Short documents, heavily formulaic writing, quoted or templated material, and writing that has been proofread hard all tend to score as more predictable than freely composed prose, which is the one property every detector is actually measuring. A free burstiness checker measures that same variation directly, which is a faster way to see the pattern than reading a description of it.

Some genres are formulaic by design. A methods section, a lab report, a standard cover letter and a literature review all reward the conventional phrase over the inventive one, because the convention is what makes the document usable to its reader. Researchers tested this directly on real writing: they ran 14,400 journal abstracts published between 1980 and 2023, years before any large language model existed, through a publicly available detector. Up to 8 percent were flagged as AI-generated regardless of publication year, and the New England Journal of Medicine's abstracts were flagged at 10.24 percent, the highest false positive rate of the five journals tested.

The test couldn't have been detecting anything about ChatGPT, since ChatGPT didn't exist for most of the years it sampled. It detected genre. More formally, as 2026's statistical analysis of the detection problem would put it, there's no way to build a detector that will distinguish between a sentence that's predictable because it was written by software and a sentence that's predictable because it's the kind of thing you write in its genre. They look exactly the same from the outside.

Heavy editing pushes in a similar direction. A first draft carries a writer's specific irregularities: the odd word order, the sentence that runs long because the thought did. A thorough proofreading pass, especially one run through another tool, tends to smooth exactly those irregularities away. The largest independent comparison of detection tools to date tested 14 of them and found that accuracy fell further, for every tool, once the sample text had been paraphrased or machine-translated.

Writing traitWhy it lowers a document's apparent unpredictabilityWhat the evidence shows
Short submissionsLess text leaves less room for the natural variation that separates human rhythm from machine rhythmScores on short documents swing more on the strength of a few unusual sentences
Formulaic or discipline-specific writingGenre convention rewards the expected phrase over the inventive one14,400 pre-ChatGPT journal abstracts (1980 to 2023) still triggered up to 8% false positives; one journal's abstracts hit 10.24%
Heavily edited or paraphrased draftsA polish pass smooths out the individual irregularities a detector reads as humanAn independent 14-tool comparison found accuracy fell further after paraphrasing or machine translation
Quoted or templated materialThe quoted passage carries its source's statistical signature, not the citing writer'sDetectors that do not exclude quotations score the surrounding document on borrowed text

Why do AI detectors misclassify non-native English speakers?

Badly, and by a wide margin. That's the conclusion of the study that first put a number on it. Stanford researchers ran seven widely used GPT detectors against 91 TOEFL essays by non-native English speakers and 88 essays by US eighth-graders. The detectors read the eighth-grade essays correctly almost every time. When it came to the TOEFL essays, which were written by adults who were fluent enough to take a graduate-level English test, the detectors averaged a false positive rate of 61.3 percent. Eighty-nine out of the 91 essays were flagged as having been generated by GPT by at least one of the seven detectors; 18 of those essays were flagged by all seven detectors.

The mechanism is the same one that governs the rest of detection: predictable vocabulary and simpler sentence construction read as low perplexity, and low perplexity reads as machine-made, regardless of who is holding the pen. Writers working in a second language lean on well-established phrasing because that is what fluency in an additional language actually looks like, and that is precisely the profile a detector was built to flag.

Writers in that position are not well served by advice to simply write more naturally, since natural is exactly what got flagged. TextPulse's page for ESL academic writers covers the phrasing patterns behind that risk in more depth, including what changes them without changing what the sentence means.

This was not an isolated case. In December 2025, a new benchmark was released, explicitly designed to evaluate bias in these detectors. It included over 200,000 samples across multiple demographics and languages. Despite being two years removed from Stanford's first study and including many different model versions of the same detectors, this benchmark revealed that the bias toward English language learners is still present.

This isn't the case for all vendors. Pangram's own technical documentation states that its classifier is not biased against non-native English speakers. The Chicago Booth researchers didn't independently test this claim but it does show the newer generation of detectors is being built with the problem in mind rather than ignoring it.

If you want to see where your own writing sits before anyone else scores it, a free perplexity checker gives you the same underlying number a detector would compute, without submitting anything to an institution first.

Who else gets flagged, and why

Non-native English speakers carry the largest documented risk, but the same statistical logic catches other writing too. Weber-Wulff's team found that all 14 tools they tested were less accurate when they ran them on paraphrased texts; this describes a meaningful share of academic writing that has been through a proofreader, an editor, or a supervisor's red pen before anyone runs a detector on it.

None of this means the number is meaningless, only that it needs corroboration before it means anything about a specific person. If you have a text that reads as uniform, similar sentence lengths, similar rhythm throughout, it will score as more machine-like regardless of who wrote it or why. This can be checked directly using a free burstiness checker instead of guessed at.

A second, less discussed group faces a related problem. Researchers compared roughly 60,000 Reddit posts, split between a subset identified as likely written by autistic authors and a general sample, and ran both through a GPT-2-based detector. Both groups stayed under 2 percent flagged overall, but the likely-autistic subset was flagged at a significantly higher rate than the general sample. Writing that leans toward literal, repetitive, tightly patterned phrasing, a documented feature of some autistic communication styles, produces exactly the low-variation text a detector is built to catch.

Why a Confident Score Is Not a Confident Accusation

False positive rate sounds like it refers to the chance of one student flagged as a problem being innocent. But no. It describes how a test works for everybody taking it. And those are different questions with different answers, just like with any screening test in medicine or security.

A March 2026 statistical analysis from a Griffith University researcher makes the underlying mathematics explicit. Treat AI detection as a test applied to a population of student writers rather than to one isolated document, and a trade-off appears: whenever some share of that population writes in a way that naturally overlaps with typical AI output, close to genre convention, close to a second-language learner's range, any detector with real power to catch AI-written work must misfire on some of that overlapping share. This is not a claim about a specific detector being poorly built. It is a property of testing a diverse population with a single document and no other evidence.

The paper's own illustrative example shows the scale involved. Suppose 10 percent of a student population writes close enough to typical AI output, and a detector is tuned to catch 80 percent of actual AI-written work. The mathematics then require an average false positive rate of at least 7.5 percent across that population. That is not a flaw in the detector's engineering: it is a direct consequence of how much that group's writing statistically overlaps with the thing being detected.

Scaled to a university of 10,000 students, that lower bound alone implies roughly 750 false flags across the population, a mathematical consequence of testing a diverse population against a single document with no other evidence to weigh. The paper's author is explicit that these particular figures are illustrative rather than measured at a real institution. The example demonstrates the shape of the problem rather than a specific prediction for any one campus.

What a defensible AI detection policy looks like

Yale already had a policy that reportedly restricts tools like GPTZero, for close to the reasons this article has laid out: a documented false positive rate and a documented bias against non-native English writers. So this isn't so much a story of one unlucky student as it is of an existing rule not being followed. When a policy actually enforces the difference between a score and a verdict, then a score becomes one input rather than a verdict.

  • Treat the score as one input, never the sole basis for a misconduct finding
  • Disclose the tool's published false positive rate to students before using it against them
  • Apply extra caution to short submissions and non-native English writing, both documented weak points
  • Guarantee a path to appeal that does not require disproving a statistic

None of this is exotic. It is closer to how any single piece of forensic evidence should be treated anywhere else: useful, checkable, and never sufficient by itself.

For an individual writer, the same principle applies before a score exists rather than after. A single number, from any tool, is an estimate rather than a verdict, and treating it as certainty in either direction, safe or caught, asks more of the number than it can support.

TextPulse's own AI humanizer tool reports a Human Score on that same logic: an estimate computed from your own text, not a detector's verdict on it, useful as a signal of how a draft currently reads rather than a promise about what any specific detector will say.

The accuracy question splits into narrower ones, each with a page of its own. The specific evidence on Turnitin's false positive rate and on ZeroGPT's accuracy is covered separately. If a score has already been used against you, there are practical pieces on what to do when you are accused of using AI, on the evidence you can assemble to prove you did not use AI, and on writing an appeal letter that answers the score on its own terms.

The number on any report will keep moving as vendors retrain their models and researchers keep testing them against harder cases. But what should not move is the habit underneath it: read a score as one measurement among several, ask what population it was validated on, and ask what the vendor is not saying about the cases it still gets wrong.

Related research: the findings above are examined at scale in Do AI Detectors Agree? An Inter-Rater Reliability Study of Nine Commercial Detectors and Human versus AI Text Classification from Stylometric Features Across 121,092 Academic Texts, TextPulse Research working papers with open data, code and a citable DOI.

Frequently Asked Questions

Yes. Weber-Wulff and colleagues found AI-detection tools 'neither accurate nor reliable' in a 2023 comparison, and non-native English writers face the highest documented risk, with one study finding an average false positive rate above 60 percent on TOEFL essays. A single score is not proof of anything by itself, which is why it should never be the only evidence behind an accusation.

Moe

PhD in natural language processing, with years spent building NLP applications end to end. Moe works on text analysis: lexical and syntactic structure, and what separates machine-generated prose from human prose statistically. He has been experimenting with computational linguistics since the early days of NLTK, spaCy and WordNet, and still writes most of his tooling in Python.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.