AI Detection

Can Turnitin Detect ChatGPT? What It Catches and What It Misses

Turnitin's AI writing report and its plagiarism report are different products with different audiences and different accuracy records. Here is what the AI score actually measures, where independent testing shows it breaks down, and why your report might not look like your roommate's.

8 min read
Illustration answering can Turnitin detect ChatGPT, showing the AI writing report displayed separately from the similarity report

Yes, but only for a specific kind of AI use, and inside a report most students never see. Turnitin's AI writing indicator flags prose that reads as statistically typical of a language model, and it does so on a growing share of the essays it scans: about 15 percent showed 80 percent or more AI-generated writing between October 2025 and February 2026, up from about 3 percent when the tool launched in April 2023. Whether it flags your specific submission depends on details that most answers to can Turnitin detect ChatGPT skip past: how the AI report differs from the plagiarism report, who actually gets shown the number, and how often that number turns out to be wrong.

This post covers the mechanics of that one product rather than academic integrity in general. It sets out what the AI writing report measures, what independent testing has found it gets right and gets wrong, and why two students submitting identical work could get different outcomes depending on which university they attend. None of that changes what your own course handbook says about acceptable AI use. It replaces a guess about a black box with a description of what is actually inside it.

Can Turnitin Detect ChatGPT? What the Report Actually Shows

The Turnitin AI writing report is a separate product from the similarity report that checks for plagiarism. They run on different models and are displayed on different tabs in the same interface. The similarity score looks at your submission and compares it to a database of previously published sources and other student papers to return the percentage that matches. The AI score does something else entirely: it's a classifier that reads your prose sentence by sentence, assigning a likelihood of being AI-generated to each sentence and then rolling those sentence-level scores up into a single document percentage. Because they measure different things, a submission can have a similarity score of 2 percent and an AI score of 60 percent, or vice versa.

This percentage is actually narrower than it sounds. This is the proportion of eligible prose that the model predicts was written by AI (not the entire document), because they remove things like code blocks, bullet points, headings and short-form writing when calculating the score. The detector can only be applied to documents that have at least 300 words of long-form prose, and as of this writing it's just for English, Spanish and Japanese text. As you might expect, set next to how AI detectors compute that score in general, the pattern here's familiar: a document-level number that looks precise is built from a narrower slice of the text than the headline figure implies.

What Shows Up as AI: ChatGPT and the Rest

Turnitin says its classifier looks for a writing pattern rather than a brand name, and its coverage has expanded repeatedly since the original GPT-3.5 and GPT-4 launch in 2023. Current guidance names the GPT-5 series, Gemini and Claude among the systems it is trained against, with new model releases added on an ongoing basis. That includes text that never touched ChatGPT specifically. A student who used Gemini, Claude or DeepSeek instead is not working around the detector simply by choosing a different company; Turnitin's own description of the model is explicit that it does not key on which vendor produced the text, only on how the text reads.

Paraphrased AI text gets its own category rather than a free pass. Turnitin's documentation separates text it judges to be AI-generated only from text it judges to be AI-generated and then AI-paraphrased, and in August 2025 it added detection aimed specifically at AI bypasser tools built to rewrite machine output past a detector. That does not mean paraphrasing is pointless or that every rewritten passage gets caught. It means the assumption that paraphrasing is invisible to Turnitin by default has been out of date for some time.

How Accurate Is Turnitin's AI Score?

Turnitin claims that it's one headline number that describes its accuracy: It says it maintains a false positive rate below 1 percent on documents where more than 20 percent of the text is flagged. The company also says it tests each new version of its models against 700,000 academic papers written before ChatGPT existed to verify this accuracy. This is a specific, testable claim about a particular set of circumstances. Take it as precisely that. Turnitin has made other claims about accuracy in 2023, shortly after its launch. The company's chief product officer acknowledged publicly that lab testing didn't predict how many false positives would show up in real-world use (especially in short submissions) and admitted the percentage of false positives was higher in the wild than in testing. He reported the sentence-level false positive rate as around 4 percent and noted Turnitin increased the minimum length of submissions that can be scored from 150 to 300 words as a result. The current minimum is still 300 words.

Independent testing adds detail the vendor figure does not. Temple University's Center for the Advancement of Teaching ran one of the more careful examinations available, in the same spirit as the independent testing of AI detectors that keeps turning up gaps between lab conditions and real submissions. Researchers submitted 120 writing samples, split evenly between fully human writing, fully AI-generated writing, AI-generated writing run through a paraphrasing tool, and hybrid texts combining human and AI contributions, and recorded what Turnitin returned for each. The tool correctly identified 93 percent of the purely human samples and 77 percent of the purely AI-generated ones.

On the tougher questions, accuracy fell even further. Turnitin accurately identified just 63 percent of the paraphrased AI samples as AI-generated and only 43 percent of the hybrid samples as neither fully human nor fully AI, which are the categories closest to how AI-helped writing actually gets produced. Hybrid text, in particular, was of interest to the researchers because they believe it will likely account for a big part (possibly the majority) of real submissions going forward as more courses assign tasks asking students to use AI for part of the work.

What was testedTurnitin's result
100% human-written samples93% correctly identified as human
100% AI-generated samples77% correctly identified as AI
AI text run through a paraphrasing tool63% correctly identified as AI
Hybrid samples, part human and part AI43% correctly identified as neither 0% nor 100%
Documents with over 20% flagged (Turnitin's own figure)Under 1% false positive rate

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Where It Breaks Down Even When the Number Looks Right

The document-level score isn't the only thing an instructor can see, and the other view is less reliable. Some interfaces let an instructor open a flag report that highlights specific sentences as likely AI-written. Temple's researchers checked those highlights against which sentences in their hybrid samples were actually AI-written, and found no relationship between the two. If a professor forwards you a screenshot of a highlighted paragraph rather than the summary percentage, the overall document score performed considerably better than the sentence-level highlights did. This matters.

Low scores are handled differently again, and more cautiously than they used to be. Since July 2024, Turnitin has displayed an asterisk in place of a number whenever its model estimates under 20 percent of a document is AI-written, because that range is where the tool is most likely to be wrong in either direction. A light AI-assisted edit, one paragraph rewritten with help and the rest written unaided, will often produce no visible percentage at all rather than a small one, which is worth knowing before treating the absence of a number as proof of anything.

Who Actually Sees the Number

The AI score is not part of the student-facing report by default. It sits on its own tab, visible to instructors and administrators, and a student sees it only if the instructor chooses to share the report or raise it directly. This is a real structural difference from the similarity score, which students can typically view themselves inside the same system once a submission has been processed.

This isn't only true for the feature itself but also for whether the feature exists at all for a particular course. This is an institutional decision above and beyond the instructor. At the top level of our institutions there're account holders who are able to switch on or off Turnitin's AI detection. One lecturer may not be able to turn it on if their university has turned it off and vice versa. A student from one university who submits similar work to another student from a different university might have a very different experience of this tool, for reasons that have nothing to do with what either of them wrote.

Why Some Universities Have Turned It Off

The University of Waterloo discontinued Turnitin's AI detection feature in September 2025, and its published reasoning is unusually direct. The university cited the tool's documented unreliability, its bias against students whose first language is not English, and its own internal testing, which included at least one case of entirely human-written text scored as 100 percent AI-generated. Weighed against the expense of the tool, the university concluded the costs outweighed the benefits and redirected the effort toward assessment redesign and AI literacy training instead.

Waterloo is one visible example of a wider split rather than an outlier. Universities that keep the feature switched on have generally added guardrails rather than treating the score as a verdict on its own, requiring a conversation with the student before anything formal starts and treating the number as a reason to open that conversation rather than close it. Reading up on how universities are handling AI policy more broadly makes the pattern clearer: whether a score reaches you at all depends on a decision made well above your seminar room, not on a fixed property of the technology itself.

What a High Score Does, and Does Not, Prove

A high AI score is evidence, not a verdict. Turnitin itself frames the percentage as one data point for an instructor's judgment rather than a finding of misconduct, and the accuracy figures above explain why that framing matters: even a document-level score that performs reasonably well misreads a meaningful share of heavily edited, paraphrased or hybrid writing, and the sentence-level highlights are markedly less reliable than the summary number. None of that guarantees any individual score is wrong. It means one percentage is a reason to ask questions rather than a finding to accept on sight. If you want to see where your own writing sits before anyone else scores it, the same low-variation, generic-phrasing patterns a classifier reacts to are checkable directly with a free perplexity checker.

It uses the same statistical approach. TextPulse's AI humanizer for students works on that same statistical basis, and so does its free burstiness checker for the sentence-rhythm half of the same measurement. It gives you an estimated Human Score based on your own draft, not a prediction of what Turnitin will say about it. No tool can responsibly promise that about a system it does not control. This is why we emphasize that this use is to give you a second read on your own writing.

None of this is likely to be the final version. Turnitin has rewritten its thresholds, its visibility rules and its model coverage more than once since 2023, and August 2026 will not be the last word either. What stays constant is the habit worth having regardless of which detector a given course uses: write with enough specific, uneven, genuinely yours detail that what a classifier happens to think becomes beside the point.

Frequently Asked Questions

Turnitin's AI writing indicator exists to answer exactly this. Can Turnitin detect ChatGPT? Usually, yes, when a document contains substantial AI writing: the classifier flags prose that reads as statistically typical of a language model, and Turnitin's own data shows about 15 percent of scanned essays now score as heavily AI-written. It also misses text regularly, with independent testing finding accuracy fell to 63 percent on paraphrased AI writing and 43 percent on hybrid text.

Moe

PhD in natural language processing, with years spent building NLP applications end to end. Moe works on text analysis: lexical and syntactic structure, and what separates machine-generated prose from human prose statistically. He has been experimenting with computational linguistics since the early days of NLTK, spaCy and WordNet, and still writes most of his tooling in Python.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.