AI Detection

How GPTZero Works

GPTZero made perplexity and burstiness famous, then quietly retired both. Here is what its classifier measures now, what a report actually shows, and where the tool says its own limits are.

Updated on 5 min read
Diagram showing how GPTZero works today: a sentence-level AI classifier replacing its original perplexity and burstiness method

Type "how gptzero works" into a search bar and most of what comes back still describes perplexity and burstiness, the two statistics its own founder made famous in the first weeks after ChatGPT's release. That description was accurate in January 2023. It has not been accurate since autumn of that year, when GPTZero's own documentation says the company replaced that method with something else entirely. The gap between what the tool did at launch and what it does now is wide enough that most explanations still circulating today describe a product that no longer exists.

Today, GPTZero runs a sentence-by-sentence classifier over a large mixed corpus of human-generated and AI-generated texts rather than a perplexity calculation. Today's GPTZero still provides a single headline number, but the arithmetic is different. The sentence-level highlighting beneath the headline number also has a different meaning than it did in 2023. Both versions of the story are worth knowing: the one that made the terms famous, and the one that actually runs today.

GPTZero result screen highly confident a passage is AI generated: 100% AI probability with per-sentence highlighting and an AI vocabulary tab
GPTZero on an AI-written paragraph: a document-level verdict, per-sentence highlighting of the sentences driving it, and an AI-vocab list. The post walks through all three layers.

The History: How Perplexity and Burstiness Made GPTZero Famous

On January 2, 2023, then-Princeton student Edward Tian released a public demo of GPTZero, just weeks after ChatGPT came out. His original explainer page describes how he used two statistics borrowed from language modeling to detect when a language model has written something. Perplexity is a score for how surprised a reference model was by the word choices in the text. Burstiness is GPTZero's name for how much that surprise fluctuated over a document. Those explainer pages are the reason both words entered everyday conversation about AI writing at all. Most people who can define perplexity or burstiness today learned the terms from GPTZero, not from a textbook.

The mechanism itself was straightforward to describe. Feed a passage through a reference language model, and prose the model finds highly predictable, low perplexity, reads as machine-typical, since generation tends to favor the likely next word over the surprising one. GPTZero's own retired documentation even published a rule of thumb: a perplexity score above 85 read as more likely human than machine. That threshold came from one vendor, on one model, at one point in time, and GPTZero itself no longer stands behind it.

How GPTZero Works Now: A Sentence-Level Classifier

GPTZero's support center states it plainly: "As of autumn 2023, GPTZero no longer uses perplexity and burstiness for its AI detection because we migrated to a deep-learning based architecture." Its current technology page, confirms the switch is total. It describes "an end-to-end deep learning approach, trained on text datasets from the web, education, and AI-generated from a range of LLMs", in which "a sentence-by-sentence classification model determines the probability and confidence that a text was created by AI". Perplexity and burstiness appear nowhere on that page.

Since then, it's also added a defense layer, a Paraphraser Shield, built to catch AI text that has been run through a rewriting tool to dodge detection. You can paste in text, upload Word documents, PDFs and image files, and up to fifty files at once. None of that infrastructure has anything to do with counting how often a reference model would have guessed the same word.

What a GPTZero Report Actually Shows You

A GPTZero result carries three parts: a classification of AI, human or mixed, a percentage expressing the model's probability that the document was AI-written, and a confidence label describing how reliable that particular prediction is. The label matters more than it might look. GPTZero ties "highly confident" to an error rate under 2 percent and "low confidence" to an error rate of 14 percent or higher, a wide enough range that the same headline percentage can carry very different weight from one report to the next.

Paid scans add sentence-level highlighting, and it is worth reading correctly. GPTZero describes the highlighted sentences as the ones disproportionately affecting the overall score, not a list of accusations against individual lines. A sentence can go unhighlighted and still have shaped the result, because the classifier reads the document as a whole rather than voting sentence by sentence.

Report elementWhat it tells you
ClassificationA single label: AI, human, or mixed
AI probability scoreA percentage reflecting how likely the model judges the document to be AI-written
Confidence labelRanges from highly confident, under a 2 percent error rate, to low confidence, 14 percent or higher
Sentence highlightingMarks sentences contributing most to the score on Advanced Scan, not every sentence connected to it

None of these report elements existed in their current form when the perplexity-and-burstiness explainer was still GPTZero's public description of itself. The interface changed alongside the method.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Why the Perplexity and Burstiness Description Still Circulates

GPTZero's original explainer is still online, still indexed, and still one of the clearer general descriptions of how perplexity and burstiness work as concepts. Search engines and other sites keep citing it, and some point to a widely shared 2026 review claiming the tool still runs a perplexity and burstiness layer quietly alongside its classifier. GPTZero's own current pages, the technology page and both support articles, say the opposite: the terms were retired, not merely hidden behind the scenes. A page describing a concept in the abstract is not the same as a page describing this tool's present-day method, and treating the two as interchangeable is most of why the confusion has lasted three years past its expiration date.

This matters beyond trivia. A student or colleague told to lower perplexity or spread out sentence length specifically to beat GPTZero is being handed advice aimed at a version of the product that stopped running in 2023. What moves a sentence classifier is a related but different question, closer to the broader account of how AI detectors actually work than to either statistic on its own. GPTZero has not stopped measuring them, on its own account they sit among seven indicators feeding the current model, but they stopped being the thing that decides the answer.

What GPTZero Says It Cannot Do

GPTZero's published benchmarks are about as strong as a vendor is likely to publish about its own product. The company reports "an accuracy rate of 99% when detecting AI-generated text versus human writing", says "we keep GPTZero's false positive rate at no more than 1% when evaluating AI versus human text", and records "a 96.5% accuracy rate" on submissions mixing AI and human writing, alongside 1.1 percent on TOEFL essays, a common stand-in for non-native English writing. Those numbers come from GPTZero's own evaluation, published in January 2025, with no independent replication cited on the page reporting them. The same page does decline the claim most vendors reach for, stating that "there is never going to be a 100% guarantee" and that anyone claiming otherwise probably has a flaw in their data or their evaluation method.

GPTZero's own limitations page is more specific than the marketing figures. The company states that accuracy improves as more text is submitted and drops at the sentence and paragraph level compared with the full document. It states its training data is mostly adult English prose, so text that departs from that pattern is harder to classify reliably. It states the classifier is not trained to catch AI text that has been heavily modified after generation. And it states plainly that other machine-generated or highly procedural writing, not just chatbot output, can get flagged the same way.

None of that is a reason to treat GPTZero's number as beyond question, and it is not the only detector worth understanding on its own terms. Readers curious what a classifier like this actually responds to can see a version of that pattern for themselves with a free perplexity checker or a free burstiness checker, two pieces of a fuller free tools collection covering the other statistical layers a report like this draws on.

GPTZero's two headline statistics have pages of their own: what perplexity measures in AI detection, and the token probability arithmetic that produces both numbers.

A tool willing to name its own blind spots this precisely is worth more trust than one that claims none, and the next unexplained percentage on a report is a good place to start asking which kind of tool produced it.

What Our Own Research Found

In the TextPulse Research detector agreement study, nine commercial AI detectors rated the same 90 academic texts. One was a 455-word hybrid text: a human-written opening and closing around a 168-word AI-generated middle. GPTZero scored this text 41.2% AI, while verdicts from the other tools on the same words ranged from 0% to 95.7% AI. The full paper, corpus, and per-tool score matrix are open access at TextPulse Research.

GPTZero on the study's hybrid text: highly confident it is a mix of AI and human.
GPTZero on the study's hybrid text: highly confident it is a mix of AI and human.

Related research: whether an AI model can be told to write like a person is tested in Do AI Models Speak Human?, a TextPulse Research working paper in which four flagship models were given a detailed style brief and a human example, then scored on a stylometric spectrum and on GPTZero against real journal prose.

Frequently Asked Questions

No. GPTZero's support center states the company stopped using perplexity and burstiness for detection as of autumn 2023, after switching to a deep-learning classifier. Its current technology page describes a sentence-by-sentence classification model and does not mention either term. The description still circulates because GPTZero's own 2023 explainer popularized both words before the method changed.

Moe

PhD in natural language processing, with years spent building NLP applications end to end. Moe works on text analysis: lexical and syntactic structure, and what separates machine-generated prose from human prose statistically. He has been experimenting with computational linguistics since the early days of NLTK, spaCy and WordNet, and still writes most of his tooling in Python.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.