AI Detection

Turnitin Similarity Score vs AI Score: Two Different Reports

A similarity score and an AI score answer different questions, so a paper can score low on one and high on the other without either number being wrong. What each measures, what triggers it, and what a high number does and does not prove.

5 min read
Turnitin similarity vs AI score comparison chart showing two independent report numbers

A report lands with two numbers on it: a similarity score of 3 percent and an AI score of 88 percent. The natural read is that they must be measuring the same thing, so the low number should mean the high number is probably wrong too. It does not work that way. Turnitin similarity vs AI score is a comparison between two separate systems that check two separate things, and a submission can score low on one and high on the other without either number being a mistake.

It's older than the AI one, built to catch text matching something that already exists. It's newer, added into the same Similarity Report as its own independent measurement, built to estimate whether prose reads as machine-generated regardless of whether it matches anything at all. They're completely independent and don't influence each other, according to Turnitin's own documentation. Everything else in this piece.

Turnitin Similarity vs AI Score: What Each Number Reports

Turnitin treats the two as different measurements bundled into one product rather than two views of one result. The Similarity Report compares a submission against a database of web content, academic publications, and previously submitted student papers, and returns the percentage of the submission's text that matches something already in that database. AI writing detection runs a different model entirely, added into that same report as its own independent score, one trained to judge how closely a passage's prose resembles machine-generated writing, with no reference to any external database at all. The fuller mechanics of how Turnitin's AI classifier reaches that number are covered separately; what matters here is that it answers a different question from the similarity engine sitting inside the same report.

What the Similarity Score Actually Measures

Similarity is a matching exercise, not a judgment. Turnitin's own guidance is explicit that the tool does not check for plagiarism. It checks for overlap and leaves the interpretation to the instructor reading the report. A quotation set off correctly, a methods-section phrase shared across an entire subfield, a reference list formatted the way every paper in the journal formats its reference list: all of it can register as a match, because the system counts overlapping strings of text rather than judging whether the overlap is legitimate.

That's also why a similarity score of zero proves less than it seems to. It means the submission contains no long, unbroken stretch of text found elsewhere in Turnitin's indexed sources. It says nothing about how the sentences that are there were produced, because producing a sentence that matches nothing in a database isn't the same skill as producing a sentence a language model wouldn't have predicted.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

What the AI Score Actually Measures

The AI score has no database to check against. A classifier reads the submitted prose and estimates, from patterns it learned across examples of human and AI-generated writing, how likely that prose is to have come from a model rather than a person. Nothing about that estimate depends on whether the exact sentence has ever appeared anywhere before. A sentence can be entirely unique, never typed by anyone until this submission, and still read as statistically typical of AI output if its word choices and rhythm follow the pattern a language model tends to produce.

This is the same statistical territory covered in full in the wider account of how AI detectors actually work: word-level predictability and sentence-level rhythm, rolled into one document score. Turnitin's version of that classifier is proprietary, but the underlying premise, that generated text tends to be more statistically typical than human writing, is the same one running underneath every tool in this category.

Similarity ScoreAI Score
What it measuresHow much of the submission's text matches other indexed documentsHow closely the prose resembles statistically typical AI-generated writing
What triggers a high numberLong quotations, common field-specific phrasing, reused material, or a match against a prior submissionUniform sentence rhythm and consistently predictable word choices across the document
What a high number does not proveThat the matched text was used improperly. Correctly cited material can still register as a match.That any single sentence was generated by AI. The classifier reads the document's overall pattern, not one line in isolation.

Can a Paper Have 0 Percent Similarity and a 90 Percent AI Score?

Yes, and the combination is not a glitch. The two scores answer different questions, so all four combinations of high and low are possible, and each one implies something different about how the text was produced.

  • Low similarity, low AI score: ordinary original writing that also reads with typical human variation. The common case for most genuine student work.
  • Low similarity, high AI score: original sentences, not copied from anywhere, that read as statistically predictable in the way generated text tends to. The signature of unedited AI writing nobody paraphrased from an existing source.
  • High similarity, low AI score: text matching an existing human-written source closely, without the statistical flatness a classifier associates with generation. Copied text, caught by the older tool rather than the newer one.
  • High similarity, high AI score: text matching an existing source that was itself AI-generated, for instance a previously submitted AI-written paper or a page of AI-produced text already indexed from the open web.

None of these combinations requires anything unusual on the software side. They fall out of the fact that one system checks for a match and the other checks for a pattern, and a piece of text can carry either property, both, or neither.

What a High AI Score Does, and Does Not, Prove

A high AI score is a statistical estimate, not a confession extracted from the text. It says the prose in front of the classifier resembles, in its word choices and rhythm, the kind of writing the model was trained to associate with AI generation. It does not identify a specific sentence as definitely machine-written, and it does not know anything about how the document was actually produced, only how it reads once finished. TextPulse takes the same position about its own numbers: what it reports is an estimated Human Score computed from the text in front of it, not a verdict about which tool, if any, touched the document.

The practical response to either number, a low similarity score or a high AI score, is the same: treat it as a reason to look closer, not as a finding to accept or dismiss on sight. Drafts, notes, and version history speak to how a document was actually written in a way a percentage never can, on either report. Anyone curious what a classifier is actually reacting to in their own draft can see the same word-level pattern for themselves with a free perplexity checker, well before any report gets generated.

TextPulse's free tools collection covers the other statistical layers worth checking before a draft goes anywhere near an instructor, not just the two compared here.

Once the two numbers are separated, the next question is what counts as a good Turnitin score. For the other detector students meet most often, how GPTZero works is covered on its own page.

The two numbers will keep sitting next to each other on the same report, and they will keep getting read as one number by people in a hurry. Knowing which question each one is actually answering is what turns a confusing pair of percentages into two separate, useful pieces of information.

Frequently Asked Questions

Turnitin similarity vs AI score comes down to what each one checks. The similarity score reports how much of a submission matches text already in Turnitin's database of web pages, publications, and past student papers. The AI score is a separate, independent measurement estimating how closely the prose reads as machine-generated, with no reference to any database at all. Turnitin's own documentation states the two figures do not influence each other.

Moe

PhD in natural language processing, with years spent building NLP applications end to end. Moe works on text analysis: lexical and syntactic structure, and what separates machine-generated prose from human prose statistically. He has been experimenting with computational linguistics since the early days of NLTK, spaCy and WordNet, and still writes most of his tooling in Python.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.