LLMs could theoretically write as diversely as humans, but they don't. Post-training and safety guardrails keep their text detectable, argues Bradley Emi, CTO of AI text detector Pangram, in a blog post. Systems like ChatGPT, Claude, or Gemini learn behavioral rules to avoid dangerous outputs or censor certain political statements. This sharply narrows their expressive range, an effect called "mode collapse."
In "mode collapse," a language model fixates on one preferred phrasing (orange curve) instead of covering the full range of human language (blue curve). With "mode coverage," the model spreads probability more evenly across phrasings. | Image: Pangram
So-called base models, the raw models before post-training, write with more variety, so Pangram's detection doesn't flag them, Emi says. The same goes for narrowly specialized fine-tunes trained only on Hemingway or certain subreddit texts, and for broken outputs like incoherent text. This only applies to non-watermarked AI text, though. Watermarks will likely always work, even with a base model's variety.
AI News Without the Hype – Curated by Humans
Subscribe now







