AI Detection

Pangram AI Detector Review: The Tool Researchers Keep Validating

Pangram publishes a false positive rate of roughly 1 in 10,000, and unlike most detector vendors it now has recent independent research pointing in the same direction. A peer-reviewed 2026 study from Vrije Universiteit Brussel found it was the only tool of four that produced satisfactory results on master's level papers. Here is what the studies measured, where the evidence is still thin, and what no detector score can prove.

7 min read
Pangram AI Detector Review: The Tool Researchers Keep Validating

Pangram is the AI detector that independent researchers keep singling out, which makes it worth a closer look than the usual vendor scorecard. Its maker, Pangram Labs, publishes a false positive rate of roughly 1 in 10,000 and a headline accuracy of 99.98 percent. Those are the company's own numbers, measured on its own benchmarks. What separates Pangram from most of the field is that, for once, recent outside research points in the same general direction.

Two independent evaluations landed within the past year: a peer-reviewed study from Vrije Universiteit Brussel, published by Springer in June 2026, and a University of Chicago Booth School working paper covered in the Chicago Booth Review in December 2025. Both placed Pangram clearly ahead of better-known rivals, including Turnitin, GPTZero, and Copyleaks. Neither study makes any detector's score proof of anything about one specific document, and both say so.

What follows is what Pangram claims about itself, what the two studies actually measured, where the evidence is still thin, and what this tool's rise says about the rest of the detection field.

What Pangram Is and Who Uses It

Pangram Labs is a small company founded in 2023 in Brooklyn by AI researchers who previously worked at Tesla and Google. The product checks pasted or uploaded text and returns an estimate of how much of it is AI generated, with sentence level highlighting, plus image detection and plagiarism checking on paid plans.

Its user base looks different from Turnitin's. Turnitin is sold to institutions and runs inside learning management systems; Pangram can be run by anyone with a free account, which means students, editors, journal staff, and admissions readers use it directly. It is the newcomer in this market, and it is increasingly the tool academic integrity researchers reach for when they want a comparison point, precisely because it keeps performing in tests the incumbents do not.

What the Vendor Claims

Pangram's homepage advertises 99.98 percent accuracy and states that its false positive rate, the rate at which human documents are incorrectly flagged as AI, is currently 1 in 10,000. A company blog post on false positives, written by its CTO, breaks that overall figure down by genre: 0.004 percent on academic essays, 0.001 percent on scientific abstracts, with weaker spots the company itself discloses, such as 0.05 percent on poetry and 0.23 percent on recipes. The company also claims its model still detects AI text after it has been processed by so called humanizer tools.

All of these are self-reported numbers from the vendor's own test sets, and the same post carries the vendor's own caveats: the model works best on longer text written in complete sentences, and the company advises against screening short bullet lists or highly formulaic text. Keep both halves of that in view. The claims are unusually specific and unusually low, and they are still claims until someone outside the company checks them.

Pangram homepage: an AI detector that actually works, with a 99.98% accuracy claim credited to University of Maryland and University of Chicago researchers
The homepage claim: 99.98% accuracy, credited to University of Maryland and University of Chicago researchers. The studies behind it are what this review reads.

What Independent Evidence Shows

The strongest test to date is a peer-reviewed study by Van Vlasselaer, Van Droogenbroeck, and Spruyt at Vrije Universiteit Brussel, published in June 2026 in the International Journal for Educational Integrity. The team built 160 master's level academic papers of at least 4,000 words each: 40 written by real students before 2019, 40 fully generated with GPT-4o Deep Research, 40 hybrid, and 40 AI generated then deliberately humanized. They ran all of them through Turnitin, GPTZero, Copyleaks, and Pangram.

On the fully AI generated papers, the study reports that Turnitin, GPTZero, and Copyleaks "completely failed to detect the AI-generated content." Turnitin scored every one of those 40 papers between 0 and 20 percent AI; the three incumbents' median scores all sat under 20 percent on papers that were 100 percent machine written. Pangram reached a strict accuracy of 65 percent and an inclusive accuracy of 97.5 percent on the same set, misclassifying one paper. On the humanized set, Pangram correctly identified AI content in 37 of 40 cases, a strict accuracy of 92.5 percent, while GPTZero managed 2.5 percent, Copyleaks 22.5 percent, and Turnitin 50 percent.

ToolFully AI papers (strict accuracy)Humanized AI papers (strict accuracy)Human papers correctly cleared
Pangram65%92.5%100%
Turnitin0%50%100%
GPTZero0%2.5%100%
Copyleaks0%22.5%100%

The second data point is a working paper by Brian Jabarian and Alex Imas at Chicago Booth, summarized in the Chicago Booth Review in December 2025. Testing detectors on roughly 2,000 human written passages across six mediums plus AI versions from four frontier models, they found Pangram's false positive rate was essentially zero across most decision thresholds, with accuracy that never dropped below 99.8 percent on their corpus and false negatives between 2 and 4 percent depending on the model. The researchers propose that institutions adopt a strict policy cap, for instance no more than 0.5 percent of human writing flagged, before trusting any detector operationally. Note the status difference: the Booth paper is a working paper, not yet peer reviewed, and it tested unmodified AI text rather than humanized academic papers.

False Positives and Who Gets Hurt

False positives are where detector damage concentrates, and the pattern documented across this industry is that flags land disproportionately on non-native English writers. The VUB study is worth reading on exactly this point: its 40 human papers were all written by non-native English speakers, and all four tools, Pangram included, cleared 100 percent of them. That is a genuinely encouraging result, and it is also 40 papers. A sample that size cannot confirm a rate like 0.004 percent; only the vendor's own much larger internal testing produces numbers that fine, and no outside party has replicated them at that scale.

The study's real world section shows why caution survives even good benchmarks. When the researchers ran Pangram over 1,163 actual master's theses from 2024 and 2025, it flagged 45.5 percent of them at some level, with a median estimated AI share of 30 percent among flagged cases, and no ground truth exists for any of those theses. The authors are direct: detection scores are useful initial flags and should never be sole evidence in high-stakes decisions. A student facing an accusation still needs process based ways to demonstrate authorship, whichever detector produced the number.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Pricing and Access

Pangram's site offers 20 free checks a day after signing up for an account, with paid subscriptions adding credit volume and features such as plagiarism detection. That access model matters as much as the accuracy numbers: unlike Turnitin, which a writer only encounters when an institution runs it against them, Pangram can be checked directly, so a writer can see the same class of signal a reviewer might see before anything is submitted.

Where Pangram Fits in the Detector Landscape

The VUB authors' summary of the field is blunt: "From the four AI detection tools studied here, at this moment only one produced satisfactory results." That sentence says as much about the incumbents as it does about Pangram. Turnitin, the tool most institutions actually rely on, scored the fully AI generated set at zero, and its separately documented false positive record has already forced universities to treat its scores carefully. All of these tools measure statistical properties of text, and the underlying mechanics explain both why Pangram's newer training approach can pull ahead and why every one of them remains probabilistic.

The honest verdict: Pangram is currently the best supported detector in independent testing, by a wide margin, on recent evidence. That evidence is also narrow. Two studies, one peer reviewed, one not, both from the past year, mostly English, mostly long form academic and web text. Detection is an arms race against models that change quarterly, and a tool that leads in 2026 holds that lead only until the next generation of text. A Pangram score is a better calibrated estimate than its rivals produce. It is still an estimate, not evidence of what any particular person did.

See the Full Comparison

Pangram is one tool in a crowded, uneven field. For how it stacks up against Turnitin, GPTZero, Copyleaks, ZeroGPT, and the rest on the same criteria, see the complete AI detectors compared guide.

What Our Own Research Found

In the TextPulse Research detector agreement study, nine commercial AI detectors rated the same 90 academic texts. One was a 455-word hybrid text: a human-written opening and closing around a 168-word AI-generated middle. Pangram scored this text 34% AI, while verdicts from the other tools on the same words ranged from 0% to 95.7% AI. The full paper, corpus, and per-tool score matrix are open access at TextPulse Research.

Pangram on the study's hybrid text: 34% AI, correctly located in the middle.
Pangram on the study's hybrid text: 34% AI, correctly located in the middle.

Frequently Asked Questions

Pangram's own claims are 99.98 percent accuracy and a false positive rate of about 1 in 10,000, measured on internal benchmarks. Unusually for this industry, recent independent work points the same way: a peer-reviewed June 2026 Vrije Universiteit Brussel study found it was the only tool of four to reliably detect fully AI generated and humanized master's level papers, and a Chicago Booth working paper found a false positive rate near zero on its corpus. The independent evidence is strong but recent and limited in scope.

Moe

PhD in natural language processing, with years spent building NLP applications end to end. Moe works on text analysis: lexical and syntactic structure, and what separates machine-generated prose from human prose statistically. He has been experimenting with computational linguistics since the early days of NLTK, spaCy and WordNet, and still writes most of his tooling in Python.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.