AI Humanization

The Em Dash Tell: Why ChatGPT Overuses It

What actually causes ChatGPT's habit of leaning on the em dash, what job the mark is doing in a sentence, and what a human editor reaches for instead: a comma, a colon, or a full stop.

Updated on 5 min read
Table showing what a human writer uses instead of the chatgpt em dash habit, by sentence function

Sometimes, a hiring manager will send a cover letter to a colleague with an added line: check the dashes. This long kind appears three times in two sentences, once as a comma, once as a colon, and once as a full stop, all in the same paragraph. The little habit that the chatgpt has is the one where you see the dash before you read the point. It's important to understand this habit properly before making any conclusions about what it actually means. That reflex, spotting the mark before reading the argument, is the chatgpt em dash habit in miniature, and it's worth understanding properly before deciding what it actually proves.

The em dash is the long dash, longer than a hyphen and longer than the dash used in a number range, that a writer drops in mid-sentence to do the job of a comma, a colon or a full stop. ChatGPT uses it more often than most edited prose does, though exactly how much more depends on which version of the model produced the paragraph in front of you. This piece stays narrowly on that one mark: what causes the habit, what job it is actually doing when it appears, and what a human editor swaps in instead.

What Causes the ChatGPT Em Dash Habit?

Nobody at OpenAI has published a technical explanation, but much of what we know is thanks to independent investigation; no company blog post needed. One good example is an in-depth analysis written by software engineer Sean Goedecke in October 2025. Instead of guessing, he tested the easy theories himself. He tried a model that saved tokens by using one mark instead of a comma. But a comma is exactly as short. He also tried a dialect preference picked up from human feedback workers, and checked a big sample of Nigerian English text, a dialect commonly associated with that workforce. In fact, it showed up less often there than in ordinary modern English, not more often.

The theory Goedecke's own testing left standing points at a change in training data between model versions. He found the habit barely present in GPT-3.5 output, then roughly ten times more frequent once GPT-4o shipped, which points at what changed in the training set rather than at the model architecture itself. His best guess is digitised books from the late 1800s and early 1900s, a period when the mark ran unusually high in printed English, well above where it sits in contemporary writing. Nobody outside the labs that trained these models can confirm it. It is the most carefully tested theory available, not a settled fact.

A separate line of research published in March 2026 points at a related cause worth naming: how much of a model's training text was formatted markdown rather than plain prose, on the theory that heavy exposure to that formatting style leaves its own fingerprint on punctuation choices generally. The honest summary is that several explanations are plausible, one is reasonably well tested, and none is confirmed by the one party that could actually confirm it.

What Job Is the Mark Doing in the Sentence?

Before picking a replacement, it helps to name the job the mark is doing, because a comma, a colon and a full stop are not interchangeable, and neither are their long-dash substitutes. Four jobs cover almost every case a reader is likely to run into.

What the mark is doingWhat a human writer uses insteadExample
Setting off a side comment inside a sentenceA pair of commas, or bracketsThe finding, replicated twice, held under both conditions.
Marking an abrupt turn or a sharp interruptionA full stop and a new sentenceThe finding held under both conditions. It did not survive a third.
Joining two closely related independent statementsA full stop, or a semicolon where the link mattersThe sample was small. The effect size was not.
Introducing a list, a summary or an explanationA colonThe result pointed one way: replication failed.

Every one of those sentences was one that could have used the mark and didn't. This piece is itself an argument that the mark is a machine tell, and it makes that argument by never using the character it describes, not once, in any sentence above or below.

Humanize your own paper

Transform your AI-assisted text and make it sound human, without touching important words or citations.

Get started free

Is Every Chatbot's Em Dash Habit the Same?

No, and the gap between systems is large. A same-prompt test across six chatbots in mid-2025, cited in TextPulse's wider guide to AI writing patterns, found three systems producing the mark roughly once every seventy words and one producing it roughly once every four hundred and seventy, while two others did not produce it at all in the sample. Frequency is a property of whichever system and version generated the paragraph in front of you, not a fixed trait of AI writing generally.

OpenAI shipped a partial fix in November 2025: a custom-instructions setting that, once a user turns it on, tells the model to stop. Sam Altman announced it himself, framing it as finally listening to a long-standing complaint. The setting is opt-in, not a default, so a huge amount of everyday ChatGPT output was never touched by it, and the habit continues wherever nobody bothered to flip the switch.

The mark's reputation as an AI tell is itself fairly recent and somewhat overstated. The backlash began around February 2025, according to commentary tracking it. A popular piece about the trend noted that there was never solid evidence that chatbots use the mark more than careful human writers already did. Both things can sit together: the habit is real and measurable in raw output, and the internet's confidence about spotting it from a single sentence runs well ahead of the actual evidence.

How to Check Your Own Draft for the Habit

Read a paragraph aloud and mark every place a long dash appears. One instance in a page is a stylistic choice, the kind human writers have made for centuries. Several in one paragraph, especially doing different jobs from the table above, is worth a second look. That density is unusual outside AI output and outside a handful of deliberately dash-heavy prose stylists.

TextPulse's free em dash remover checks that density automatically and proposes the comma, colon, full stop or restructure each instance calls for, which is faster than counting by hand across a long document. The same free tools collection covers the other levels a machine-drafted paragraph tends to signal at once, word choice, sentence rhythm and structure, for a check that goes beyond punctuation alone.

Filler phrases do the same work at sentence level, and there is a list of the ones worth cutting first. Removing this vocabulary from a draft you have already written takes a different approach again.

None of this required special software to check, just attention and a willingness to read the sentence again. This piece made the same choice throughout its own drafting: reach for a comma, a colon or a full stop, and never once for the mark it spent fifteen hundred words describing.

Related research: the findings above are examined at scale in Stylometric Fingerprints of AI Rewriting: Punctuation, Syntax, and Model Attribution, a TextPulse Research working paper with open data, code and a citable DOI. Whether prompting alone can make a model write like a person is tested in Do AI Models Speak Human?.

Frequently Asked Questions

The ChatGPT em dash habit comes mostly from what changed in its training data between model versions, not a deliberate design choice. Independent testing found the mark far more common in GPT-4o output than in GPT-3.5, and the leading theory points at heavy exposure to older printed books from a period when the mark ran unusually high. OpenAI has not published its own explanation, so this remains the best-supported theory rather than a confirmed fact.

Mark

Content strategist at TextPulse, here since the company started. Mark writes the product and technical coverage: how the humanizer works under the hood, what changes in each release, and what a specification actually means for your writing. His reviews of writing software come from using them on real documents rather than reading a feature list.

Stay updated on AI humanization

Get tips on academic writing, AI detection, and humanization delivered to your inbox.

No spam. Unsubscribe anytime.