Last updated:
🕵️ Stylometry and the Pseudpocalypse: How Your Writing Identifies You
Everyone has a writing fingerprint — and machine learning has made reading it nearly effortless. How authorship attribution works, why online pseudonymity may be dying, the famous unmaskings from the Unabomber to J.K. Rowling, and what actually defends your anonymity.
Key findings
Stylometry identifies authors from unconscious writing habits — function-word frequencies, sentence length, punctuation — that persist even when you disguise your identity or change topics. The landmark result: Narayanan's 2012 Internet-Scale Author Identification study identified anonymous authors from a pool of 100,000 with meaningful accuracy; smaller pools of ~100 authors reach up to 94%. The threat survives context-switching — anonymous blog posts have been linked to a person's emails on unrelated subjects. LLMs are the game-changer: they've turned a specialist task into something anyone can do by just asking, and one recent model identified a writer from the first 1,000 words of a draft. This drives the "pseudpocalypse" (a term from the writer Dynomight): the conjecture that online pseudonymity is dying, because enough text under any name can be linked to you. Famous unmaskings: the Unabomber (his brother recognized the style) and J.K. Rowling as Robert Galbraith (2013, computational analysis). The crucial caveat: stylometry proves nothing — it's "much less reliable than DNA," a probability not a certainty. Defenses exist (adversarial stylometry: obfuscation, imitation, translation) but are imperfect against a determined adversary.
You have a writing fingerprint you can't feel and can't easily hide. For decades, reading it took a specialist. Now a language model can often do it from a paragraph — which turns a niche forensic technique into a live threat to anyone who writes under a name they want to keep separate. This is how stylometry works, how far it reaches, and what genuinely defends against it — with the honest limits stated plainly.
By Ned Walsch · Last updated July 2026
Your writing fingerprint
Stylometry rests on a simple, well-established fact: everyone writes with an idiolect — a personal statistical signature built from habits you're not aware of and can't easily suppress. Not the content of what you say, but the mechanics: how often you reach for function words like "the", "of", or "while"; your average sentence length; your punctuation rhythm; the word-pairs you favor; thousands of small tics. Because these are largely unconscious, they survive deliberate disguise and topic changes.
The technique is old. Its first famous win was in 1964, when Mosteller and Wallace resolved the disputed authorship of the Federalist Papers by measuring tiny differences in common-word usage — Madison wrote "whilst" where Hamilton wrote "while", and the frequency of a bland word like "by" differed enough to assign the contested essays. What's changed is the power and reach of the tools.
How far it reaches
The pseudpocalypse
The writer Dynomight coined a memorable frame for where this leads: the "pseudpocalypse" — the conjecture that if you put any significant amount of text online under different names, those identities can eventually be linked using only the text. Online pseudonymity, in other words, may be dying.
The stronger version generalizes beyond writing. As identification technology improves, any rich channel you use to interact with the world tends to reveal you:
Why LLMs changed everything
Traditional stylometry was work: build a feature set, train a classifier, assemble a labeled corpus of candidate authors. That expertise barrier is why, for decades, it was mostly an academic and forensic tool rather than an everyday threat.
The famous unmaskings
Both cases teach the same lesson — sustained writing under a pseudonym is a persistent link to your identity — but the Rowling case also shows the technique's limit: the analysis proved the book was either Rowling or someone who wrote very like her. The tip-off did the rest.
The honest limit: it proves nothing
Can you defend against it?
Partly, with real and sustained effort. There's a research field — adversarial stylometry — devoted to exactly this, and it's found workable but imperfect defenses:
The catch is consistency: your habits reassert themselves the instant you stop concentrating, and maintaining a disguised style across a large body of text is genuinely hard. None of this is foolproof against a determined, well-resourced adversary.
If you actually need anonymity
Design around the fact that writing style is an attack surface. The most reliable defenses are behavioral:
- Minimize volume. These methods need a sizable sample — the less you write under a sensitive identity, the less there is to fingerprint.
- Keep identities genuinely separate, and remember the link survives across topics, so writing about unrelated subjects is not protection.
- Use intermediaries. For high-stakes anonymity, have someone rewrite your words, write collectively so no single style dominates, or have your text heavily edited.
- Plan for failure. Assume a sophisticated adversary might identify you, and design your operational security so that identification isn't catastrophic — stylometry is only one of many channels through which anonymity leaks. See our OpSec guide for the rest.
Frequently asked questions
What is stylometry?
Stylometry is the statistical analysis of writing style to identify or characterize an author. The core assumption is that everyone has an idiolect — a personal fingerprint in how they write, made up of habits they're not consciously aware of and can't easily suppress: how often they use function words like 'the', 'of' and 'while', their average sentence length, punctuation patterns, favored word pairs, and thousands of other subtle features. Because these habits are largely unconscious, they persist even when a writer is deliberately trying to disguise their identity or is writing about a completely different subject. The field is also called authorship attribution, and it's old — its first famous success was a 1964 study by Mosteller and Wallace that settled the disputed authorship of the Federalist Papers by analyzing tiny differences in common-word usage, such as the fact that Madison wrote 'whilst' where Hamilton wrote 'while'. What's changed recently is that machine learning, and now large language models, have made it vastly more powerful and accessible.
Can writing style really identify me?
Given enough text and a defined pool of suspects, yes — often with unsettling accuracy, and the research bears this out. The landmark demonstration is a 2012 study by Arvind Narayanan and colleagues titled 'On the Feasibility of Internet-Scale Author Identification', which showed that stylometry could identify the author of an anonymous piece of text from a pool of 100,000 possible authors with meaningful accuracy — the correct author was the top match around 20 percent of the time, and in the top few matches far more often. That's at internet scale, from a huge suspect pool. With smaller pools the accuracy climbs steeply: studies differentiating around 100 authors have reported accuracy as high as 94 percent. And crucially, the identification survives context-switching — earlier work by Rao and Rohatgi showed that anonymous posts could be linked to a person's known writing even when the two samples were about entirely different topics, and the threat persists across domains, such as linking blog posts to emails. Your writing style is, in a real sense, a biometric.
What is the 'pseudpocalypse'?
It's a term popularized by the writer Dynomight for the idea that online pseudonymity is dying — the conjecture that if you put any significant amount of text on the internet under different names, those identities can eventually be linked using only the text. The stronger version is a 'generalized pseudpocalypse': the notion that as identification technology improves, interacting with the world through almost any rich channel will end up revealing who you are. Mask your face and use a voice changer, and you can still be identified by gait or body shape; ban license plates, and cars can be tracked by unique engine sounds or tire wear; lock down your browser and pay only in Monero through three VPNs, and you might still be identified by how you move the mouse or scroll. The writing-style version is the most immediately real, because the tools already exist and are improving fast. Notably, the author observed that a current large language model, given the first thousand words of a draft, correctly identified him — the kind of near-effortless attribution that used to require a specialist and a research paper.
How have LLMs changed stylometry?
They've collapsed the cost and expertise required, which is the whole reason this shifted from an academic curiosity to a practical privacy threat. Traditional stylometry required building feature sets, training classifiers, and assembling a labeled corpus of candidate authors — real work for a specialist. Large language models can now often guess authorship, or infer an author's traits, from a short sample with no special setup at all, simply by being asked. This 'democratization' is the dangerous part: it takes a previously labor-intensive, expert task and turns it into something automated, streamlined and widely accessible. A capability that a well-resourced adversary once had is now available to almost anyone. And the models keep improving at identifying increasingly obscure writers from increasingly small amounts of text, including across different registers and subjects. The trajectory points toward a world where any substantial body of pseudonymous writing can be linked to its author with minimal effort.
Who has been unmasked by their writing?
The two most famous cases bookend the technology's history. The Unabomber, Ted Kaczynski, was a domestic terrorist whose identity remained unknown through a long bombing campaign — until his manifesto was published and his brother recognized the writing style and phrasing, prompting the investigation that caught him. That was human recognition of idiolect, before automation. The landmark computational case is J.K. Rowling: in 2013 she published the detective novel 'The Cuckoo's Calling' under the pen name Robert Galbraith, and after a tip-off, computational linguists including Patrick Juola compared the book's common-word and character-n-gram patterns against Rowling's known work and found a strong match. It was enough for the Sunday Times to confront her agent, and she confirmed it. It's worth stressing what that case did and didn't prove — the analysis established that the book was either by Rowling or by someone who wrote very much like her, not a certainty. But combined with the tip, it was decisive. Both cases illustrate the same lesson: sustained writing under a pseudonym is a persistent link to your identity.
Does stylometry actually 'prove' who wrote something?
No, and this is the most important caveat, one that even enthusiasts insist on. Stylometry is probabilistic, not conclusive. As the computational linguist who worked the Rowling case and others have emphasized, it's far less reliable than something like DNA — a stylometric match tells you the text was likely written by a particular person or by someone with a very similar style, not that it definitively was. Writing style varies within a single author across genres, moods, and years, which is exactly why identification is hard and why results come with error rates rather than certainties. Historically this unreliability, and some high-profile forensic failures, limited stylometry's admissibility as courtroom evidence, though rigorous analyses have gained more acceptance over time. The practical upshot: stylometry is a powerful tool for generating leads and raising or lowering probabilities, and it can absolutely be enough to break anonymity when combined with other clues — but treat any single stylometric 'identification' as strong suspicion, not proof.
Can I defend my anonymity against stylometry?
Partly, with real effort, and the honest answer is that it's hard and imperfect. There's a research field called adversarial stylometry devoted to exactly this, and it has found some workable defenses. The main approaches are: obfuscation (consciously changing your writing to push its statistical features toward the average, making them less distinctive), imitation (deliberately writing in the style of someone else), and translation (round-tripping text through machine translation to launder stylistic markers, though this degrades quality). Studies have shown that even untrained people, given basic guidance, can reduce an attacker's accuracy substantially — in some cases down toward random guessing — but doing it consistently across a large body of text is genuinely difficult, because your habits reassert themselves the moment you stop concentrating. There are also tools designed to help, historically things like Anonymouth, which flag the features making your writing identifiable. None of this is foolproof against a determined, well-resourced adversary.
If I want to be genuinely anonymous, what should I actually do?
Accept that writing style is an attack surface, and design around it rather than assuming a pseudonym protects you. The most reliable defenses are behavioral, not technical. Minimize the volume of text you publish under any identity you want to keep separate — the less you write, the less signal there is to fingerprint, because these methods need a reasonably large sample to work. Keep pseudonymous writing genuinely separate from your real-name writing, and be aware that the link survives even across different topics, so writing about unrelated subjects is not protection. For genuinely high-stakes anonymity — whistleblowing, dissent, sensitive reporting — consider using an intermediary who rewrites your words, writing collectively so no single style dominates, having your text heavily edited, or simply accepting that a sophisticated adversary may be able to identify you and planning accordingly. And combine this with the rest of a real operational-security posture, because stylometry is only one of many channels through which anonymity leaks. The blunt truth of the pseudpocalypse argument is that perfect textual anonymity, at scale, may no longer be achievable — so the goal is raising the cost of attribution, not guaranteeing it's impossible.