With synthetic text becoming commonplace across the internet, readers and linguists have grown increasingly alert to subtle stylistic signatures that distinguish software output from genuine human thought. While earlier dead giveaways such as the overuse of em-dashes and the frequent appearance of words like delve have largely been engineered out of cutting-edge models, researchers have discovered that modern frontier systems continue to rely on repetitive linguistic habits when drafting prose.
Graphite Study Identifies Thirteen Thousand Statistical Tells
A rigorous examination conducted by marketing organization Graphite evaluated the prose tendencies of leading artificial intelligence engines to map each architecture's recurring vocabulary and stylistic choices. Even though punctuation hallmarks like heavy em-dash deployments have been sharply curtailed, models continue to lean into contrast-heavy sentence designs, with every release cycle presenting its own peculiar habits. What stood out most to the investigative team was the sheer breadth of these linguistic patterns. Graphite cataloged 13,000 distinct phrases that occurred at least twice as frequently in synthetic prose compared to human-authored material, setting that double-frequency threshold as their definition of a diagnostic tell.
Greg Druck, chief AI officer at Graphite, observed that Claude models developed by Anthropic have actually been moving closer to natural human word distributions with successive updates. By contrast, OpenAI's GPT systems appear to be drifting further away from those human baselines over time, showing a divergence in how these labs tune their outputs.
Methodology Built on Ten Thousand Pre-Launch Articles
Evaluating machine-generated phrasing at scale required an objective and structured testing framework. Graphite assembled a benchmark library consisting of 10,000 articles published prior to the debut of ChatGPT, using this material as an unpolluted baseline of authentic human writing. Researchers subsequently tasked multiple frontier systems with rewriting these pieces using only short summaries, deliberately neutralizing any direct textual influence from the original phrasing. With parallel sets of articles produced by both human writers and individual machine models, the team was able to measure the exact frequency of specific words and phrases while tracking broad architectural traits in sentence structure.
Distinct Habits in Claude Opus 5.5 and OpenAI Astra
Graphite's findings show that Claude Opus 5.5 exhibits a strong preference for the descriptor dependable, applying it 23 times more often than human writers do in comparable contexts. While Opus 5.5 has largely abandoned the predictable construction stating that something is not X, it is Y, it frequently resorts to a modified variant, claiming that a concept is more than an X, it is a Y.
More prominently, Opus 5.5 displays a persistent compulsion to instruct the audience on significance. The specific wording this matters turned up 116 times more frequently in Opus text than in human writing, while the explanatory framing why X matters appeared 92 times more often than normal.
OpenAI's Astra model relies on an entirely different set of recurring habits. It shows an affinity for describing another dimension of various subjects and repeatedly softens assertions by stating that an intervention may provide or can provide a particular outcome. Astra's most dominant structural feature is what Graphite terms corrective framing, wherein an idea is described as not simply X or framed through an alternative phrase like rather than relying on X. According to the research, such corrective clauses appeared more than 100 times more often in text produced by Astra than in human writing.
Punctuation Fixes Fail to Reduce Overall Quirk Count
Frontier development teams have clearly taken steps to address public criticism surrounding punctuation, specifically the historical overuse of em-dashes. Within Graphite's dataset, Opus 5.5 reduced its use of the mark by 99 percent compared to Opus 5. Astra deployed the punctuation 88 percent less frequently than the human benchmark group, while Gemini 3.1 Pro virtually eliminated the mark from its generated text altogether.
Yet, while individual tics are suppressed, Graphite determined that the aggregate volume of tells remains virtually unchanged. Druck noted that these diagnostic flags are not diminishing overall. While the most notorious markers are successfully removed during fine-tuning, alternative patterns emerge to take their place, leaving each iteration with its own fingerprint.
Parameter Scale Limits Developer Control Over Phrasing
The persistence of these verbal signatures remains striking given how aggressively engineering teams advertise their systems' conversational fluency. When Anthropic launched Opus 5.5, the company highlighted that the architecture communicated in a noticeably more organic fashion than earlier versions, adding that testing feedback praised the text for being more lucid and effortless to digest. Similarly, OpenAI made comparable claims while introducing the Sol and Luna editions of GPT-6, telling users to anticipate improved clarity, diminished corporate jargon, and fewer awkward phrasing habits.
Druck expressed doubt that research teams will ever be able to purge these stylistic artifacts completely. He suggested that commercial labs may have less granular authority over subtle language nuances than people assume. With models operating on billions of parameters, engineering teams can only execute a finite number of evaluations, allowing recurring phrasing patterns to slip through unnoticed into production builds.



















