AI Humanization

AI Writing Patterns: The Words, Rhythm and Punctuation That Give It Away

Machine-written prose has a signature: predictable words, uniform sentence rhythm, a dash doing the work of a comma, and an argument shape that restates its own heading. Here is what each level actually looks like, with sourced examples instead of a vibe.

9 min read
Table of AI writing patterns, overused words such as "delve" and "underscore", next to plainer human alternatives

Read this sentence once and see if you already know where it came from: researchers continue to highlight the complex interaction between motivation and achievement, and a growing body of work highlights the key role that self-regulation plays in academic success. Nothing about it is wrong. Every clause is grammatical, every word is a real word, and a marker wouldn't dock a single point. But it still reads like it fell out of a machine, because it did. This is what this piece is about.

AI writing patterns are the repeatable habits, in word choice, sentence rhythm, punctuation and structure, that make machine-drafted prose recognisable on sight. They cluster at four levels: the words a model reaches for, the shape of its sentences, the punctuation it leans on, and the way it structures an argument as a whole. None of them alone proves anything. Together, inside one paragraph, they are difficult to miss.

The machine that does this is called a statistical detector. It scores how predictable each word is, according to what a language model expects. A report contains the perplexity and burstiness math underneath that score, covered in more detail in another piece on how AI detectors work. The rest of this article isn't about a model, but rather it is about a catalogue of AI writing, a reader can hear, count or point to on the page. This can matter, even if a detector is wrong, unavailable, or beside the point. What follows here needs no model and no subscription.

Why Do AI Writing Patterns Sound Robotic?

A language model is doing something different from a person at the sentence level: picking the statistically likely next word, over and over, inside a system trained to sound competent rather than distinctive. Competence at that scale produces smooth, correct, forgettable prose, which is the root of the robotic feeling. A reader often cannot say what is wrong with a paragraph like that, only that it reads flat, and the four levels below are where the flatness lives.

Word choice is the level readers notice first: formal, theatrical words showing up where they do not belong. Sentence rhythm is quieter, a paragraph where every sentence runs to roughly the same length. Punctuation and formatting carry their own signature, a dash used as a universal joint, a bullet list breaking up text that was never a list in the writer's head. Structure is the least visible, a piece that restates its own heading, covers exactly three points, then closes by summarising itself.

Start with the words: the easiest to point to, and the best documented.

What Words Does AI Overuse?

You see these formal, slightly theatrical words again and again: "delve", "underscore", "showcase" and "pivotal" are among the most documented. Dmitry Kobak and three co-authors set out to identify the language. They analyzed 15.1 million PubMed abstracts from 2010 to 2024 and identified a group of words that saw a sharp increase in frequency when ChatGPT was released. "Delve" went the farthest; the study, which was published in Science Advances in July 2025, estimated its frequency in 2024 as about 28 times higher than before 2023. The next closest were "underscore" and "showcase". The authors estimate at least 13.5 percent of 2024 abstracts carry signs of AI assistance, climbing above 40 percent in some countries and publishers.

Why models converge on these particular words is less settled than the pattern itself. A COLING 2025 study tested twenty-one overrepresented words against model architecture, training data and the human feedback used to fine-tune chatbots, and found the feedback stage the most plausible contributor without fully explaining the effect. The interpretation is still argued over, too: a 2026 exchange in the journal PNAS pushed back on how directly a word-frequency shift measures AI use, worth remembering before treating any single figure here as settled fact.

That's not just true for journal abstracts. In a study comparing half a million student and AI-generated essays, researchers compared the papers' length and word choice and found the human essays ran longer overall and used a wider vocabulary. In a study of peer reviews submitted to four machine learning conferences, including ICLR 2024 and NeurIPS 2023, researchers used the same technique to estimate that between six and fifteen percent of the review text was written considerably by an LLM. It was more likely to be affected by an LLM when it was filed near the deadline.

Word or phrase AI overusesWhat a careful writer uses instead
delve (into)look at, examine
underscoreshow, stress
showcasepresent, demonstrate
pivotalmain, deciding
robustsolid, reliable
intricatecomplex, detailed
meticulouscareful, thorough
tapestrymix, range
testamentsign, evidence
boastshas, includes
fosterencourage, support
garnergain, attract
multifacetedmany-sided, complex
holisticwhole, overall
nuancedcareful, precise
invaluableuseful, valuable
realmfield, area
embark onbegin, start
unlockopen up, make possible
interplayrelationship, link

None of these words is wrong on its own. A methods section calling a finding robust is using the word correctly, and a historian describing an intricate legal argument is not confessing to a chatbot. What reads as machine-made is density: three or four of these words in one paragraph, doing no specific work, stacked the way a model stacks them because each was independently the safest next choice. A quick pass with TextPulse's free AI word cleaner flags that density directly.

Why Do AI Sentences All Sound the Same Length?

A model builds each sentence the way it built the last one, rather than compressing a point into six words in one place and unspooling it across thirty somewhere else, so the output tends to sit at nearly the same length throughout. That variation, or the lack of it, is something you can hear by reading a paragraph aloud, a different exercise from the burstiness score a detector calculates from the same underlying property.

Two smaller habits give the rhythm away up close. A model likes to close a sentence with a dangling participle that gestures at significance without adding information, a claim followed by highlighting the need for further research or reflecting a broader shift in the field. It also reaches for a paired structure to sound balanced, along the lines of the method is not only efficient but also scalable. Said once, that construction is unremarkable. Said in every paragraph, it sounds like a template with the nouns swapped out.

The Rule of Three

At the phrase level too, groups of three are the model's default rhythm: three adjectives before a noun, three examples in a list, three clauses built the same way in a row. Occasional groups of three have always been part of English. A well-placed group of three isn't a symptom of anything. The tell is frequency. Three isn't the problem. Three showing up in nearly every paragraph is.

Counting is enough to catch most of this without any special training. Read a paragraph and mark the length of each sentence in words. If all your marks are within a few words of the last one, the rhythm is the tell. Use TextPulse's free burstiness checker to automate counting and chart it, useful for a passage too long to eyeball one sentence at a time.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

The Em Dash and Other Punctuation Tells

Why the Mark Became the Go-To

The em dash is probably the most blamed punctuation mark in machine writing. You might know it as the long dash that replaces a comma, a colon or a full stop to connect two halves of a sentence. The em dash isn't new and it's not something invented by AI. For as long as style guides have been around, from Chicago to AP, they've allowed you to use it for emphasis and interruption. And lots of working writers use it intentionally.

The likely reason it shows up so often in machine output is mundane: the mark is common in the edited, professionally set prose language models are trained on, and it lets a sentence keep moving without committing to a full stop. A model built to produce fluent, connected text reaches for that flexibility often, more often than most people do when drafting by hand.

How often varies enormously by which system produced the text. A side-by-side test run in mid-2025 gave six chatbots the same prompt and counted the marks: three systems returned it eight or nine times in under six hundred words, one used it twice in nearly a thousand words, and two did not use it at all. Frequency is a property of whichever model produced the paragraph in front of you, not a fixed trait of AI writing. OpenAI added a setting in November 2025 that lets a user tell ChatGPT to stop using the mark and actually comply, which makes even this tell partly optional now.

None of that makes the mark itself evidence of anything. Plenty of careful writers use it deliberately, and plenty of student essays that never touched a chatbot use it too, sometimes because a teacher taught them to. What is worth checking is density: a page leaning on the mark every second sentence reads differently from one that uses it twice on purpose. TextPulse's free em dash remover checks that density and proposes the comma, colon, full stop or hyphen each instance calls for.

Bullets, Bold Text and Headers

A list in the middle of a paragraph because you wanted to separate out your thoughts into a bulleted list even though there was no such thing as a bulleted list in your mind. Headings in title case when the rest of the document is in sentence case. A bolded phrase at the beginning of each bullet point like it's a mini headline. The same level covers formatting decisions below the sentence: None of these are punctuation marks, but they all come from the same place: a model formatting for a chat window, where visual structure substitutes for the transitions prose would normally carry.

Structure and Argument Shape

The last level sits above the sentence and is the hardest to fix with a word swap, because it is the shape of the argument itself. A machine-written section often opens by restating its own heading in slightly different words, moves through two or three supporting points given roughly equal weight regardless of how much each one actually matters, then closes by summarising what it just said rather than adding anything new.

A related habit shows up constantly in longer pieces: a section lists what something does well, pivots on a word such as however, lists challenges in the same tidy three, then gestures at a future that will presumably resolve them. Real arguments are lumpier than that. A human writer usually cares more about one point than the other two, gives it more space, and does not owe the reader a symmetrical close.

This is also the level where formulaic academic writing can look deceptively similar to machine writing, since a methods section or a literature review is supposed to follow a fixed shape. The difference sits inside that shape. A machine-written literature review name-checks a field's usual points in the usual order. A human one argues about which points matter and why, even inside a rigid template, because an actual position is being defended.

Do These Tells Prove a Text Is AI-Written?

But no single tell proves anything, and neither does a full set of them. Under pressure from an editor or professor, a careful academic writer can create paragraphs with several of these features (such as lack of pronouns) but without the help of a chatbot. Non-native English writers may also use safe words they were taught in class, and students who are taught to use formal connectors might use many of these features too. A word list isn't proof of misconduct; it's a mistake that can get careful writers falsely accused. We should call out this mistake, not quietly repeat it.

What a tool can responsibly do is show the pattern, not adjudicate the writing. What we have to offer with TextPulse's AI humanizer is an estimate of a Human Score based on the text in front of it. This is a reading of where a draft sits against these surface patterns, not a verdict from any detector and not a promise about what a different tool says tomorrow. It's a number you can use as a starting point for a decision, the way a spellchecker flags a word without deciding whether the sentence around it is any good.

Next time you're reading your prose and something in it doesn't feel quite right, don't just find a synonym to replace it. Try reading it out loud instead. Count the number of words in each sentence. Ask yourself what exactly each word is doing there. That's the same read a careful editor was running long before any of this had a name.

Frequently Asked Questions

Because several of the surface patterns that read as machine-made, formal vocabulary, uniform sentence length, a tidy three-part structure, also show up in careful human writing done under deadline pressure or heavy editing. These are tells, not proof of origin. A paragraph can carry every one of them and still be entirely your own work.

Mark

Content strategist at TextPulse, here since the company started. Mark writes the product and technical coverage: how the humanizer works under the hood, what changes in each release, and what a specification actually means for your writing. His reviews of writing software come from using them on real documents rather than reading a feature list.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.