TextPulse Research

Open studies on AI generated text
and how machines judge it.

We study the questions our users face on a daily basis. How AI detectors behave, how human and AI writing styles differ linguistically, and what happens when the two are fused. Every study comes with a publicly available corpus, score data, and analysis code, so extended work is encouraged.

Working papers

Modelometry: Classification of Flagship AI Model Families Based on Inter-Model Stylometry

Studies that measure the output of more than one AI language model consistently find that the models write differently, in the same manner that human authors have different writing styles. We coin the term modelometry for the measurement and attribution of the inter-model writing style of AI systems, and modelolect for the style itself, formed…

Do AI Models Speak Human?

Users who want an AI model to write like a human usually assume that the machine register is a default that an instruction can lift. We tested that assumption in this study. Ten passages of human-written published academic prose gave ten topics. Four flagship models (GPT-5.6 Sol, Claude Opus 5, Gemini 3.7 Flash and DeepSeek…

A Controlled Comparison of AI Text Humanizers on Academic Writing

AI text humanizers rewrite machine-generated text so that AI detectors classify it as human. Humanizer vendors advertise pass rates, but no published study measures what the rewriting does to the output text itself. We passed 48 AI-generated academic texts (about 300 words each, four disciplines, three citation styles, four AI model generator families) through eleven…

Non-Native English Writing and the False Positives of Stylometric AI Text Detection

A widely cited study found that perplexity-based AI text detectors flag the majority of essays by non-native English writers as AI-generated while sparing native writers, and the finding has impacted the debate on AI detection in education. This study measures whether the bias is true for a transparent stylometric classifier. We score all 5,600 essays…

The Detectability of Partially AI-Rewritten Academic Documents from Stylometric Features

Real documents are often partly AI-processed, with a few sections AI-rewritten and the rest left as human-written. This study measures how a stylometric human-versus-AI binary responds as the AI-rewritten share of a document increases from 0 to 100 percent. From 25,561 content-aligned pairs of human-written academic texts and their AI rewrites, we splice documents in…

Text Length and the Reliability of Human versus AI Text Classification in Academic Writing

Our previous study classified human-written and AI-rewritten academic texts from interpretable stylometric features with an AUC of 0.936 on texts with a median length near 290 words. This study measures how that reliability depends on text length. From 25,561 text pairs in which both the human text and its AI rewrite contain at least 300…

Human versus AI Text Classification from Stylometric Features Across 121,092 Academic Texts

This study treats the separation of human-written and AI-rewritten academic text as a plain text classification task. The data are 60,306 human-written academic texts and 60,786 AI rewrites of those texts, produced by eight AI model configurations across six model families, each text described by 47 interpretable stylometric features such as word length, passive voice…

Stylometric Fingerprints of AI Rewriting: Punctuation, Syntax, and Model Attribution Across 60,786 Paired Texts

When a language model rewrites a human text, it changes more than the words. This study measures what occurs to the rest of the linguistic style. We compared 60,786 human-written academic texts with AI rewrites of the same texts produced by eight AI models across six model families, and computed 49 stylometric features for every…

Sentence-Length Burstiness as a Cross-Disciplinary and Cross-Model Signal of AI Rewriting

Burstiness, the uneven rhythm of sentence lengths in a text, is one of the most explicit differences between human and AI writing, and among the least precisely measured. This study measures it at scale. We compared 60,779 human-written academic texts with rewrites of the same texts generated by eight AI model configurations across six model…

The Vocabulary Fingerprint of AI Rewriting: Common Words AI Language Models Prioritize

We compared 60,786 human-written academic texts with AI rewrites of the same texts, produced by eight model configurations across six model families, 40.4 million tokens in total, and measured which words the models inject and suppress. The signature is consistent: formal connectives and Latinate substitutions such as "thereby" (13 times the human rate), "consequently" (11…

What Happens to Citations When AI Rewrites Academic Text? A Large-Scale Paired Audit

We audited 60,786 paired passages, each a human-written academic excerpt and a machine rewrite of that same excerpt, produced by eight model configurations across six model families. The pairs contain 213,881 in-text citation marks, and every rewrite citation can be checked exactly against its source. 96.9 percent of citation marks survived rewriting unchanged. 2.26 percent…

Do AI Models Invent References? A Verification Audit of Citations in AI-Generated Academic Text

Five current model families (DeepSeek, Mistral, OpenAI, Anthropic, and Gemini, 2026) were audited under one protocol across 30 academic topics. Writing in prose, the models embedded 194 author-year citations and none was an outright fabrication. Asked for full reference lists, the same models produced 1,500 references, of which 15.2 percent were fabricated or attributed a…

Do AI Detectors Agree? An Inter-Rater Reliability Study of Commercial AI Text Detectors on Academic Writing

Nine commercial AI text detectors (Turnitin, GPTZero, Originality.ai, Pangram, Copyleaks, ZeroGPT, Winston, Sapling, and QuillBot) were treated as independent raters of 90 academic texts: 30 purely human-written before 2022, 30 AI-generated by five model families, and 30 hybrid (human + AI) splices. Overall agreement is substantial (Krippendorff's alpha 0.71, Fleiss' kappa 0.78), but on hybrid…

Further studies in this series are ongoing.