back
86 comments
There's been some interesting threads about stylometry over the years [1]. The top link was quite decent at unmasking HN alt accounts with basic ngram analysis whipped up in one day [2].

1. https://hn.algolia.com/?q=stylometry

2. https://news.ycombinator.com/item?id=33756141

I remember this! it got me pretty good. I have a bad habit of generating alts because I forget the password.

It makes me wonder if we could use non-instruct LLMs to slightly alter the wording of text while keeping the meaning the same meaning. Perhaps by using perplexity or some other metric. I don't know, maybe you compare the distance of the "meaning" vectors.

You might also want to have some "style" vector associated with each pseudonym. For example, I might want it to produce british english under a certain pseudonym, and simulate an ESL speaker under another.

Essentially, you would want some way of re-styling text. The basic way to do this would be to run the same sylometry tools the hunter uses and manually make synonym word/phrase subsitutions to lower your similarity.

It's a cat and mouse game, but I think the mouse eventually wins. Consider a program that translates your english writing programmaticially into a low-entropy symbolic form and then translates them back to english in a procedural manner. Basicially you design an intermediary language that cannot contain style. It would be boring to read but it would remove all the style.

If that's as effective as comments suggest, then it seems indeed unmasking pseudonyms has only been a matter of effort for a while.

I still think people underestimate the power of even minor inconvenience. While you can't just click a button to reveal all pseudonyms of someone (e.g. you need to download some obscure tool or even perform statistical analysis yourself), I think this provides significant (and surprising?) protection for pseudonymous individuals, I'd say even (although to a reduced extent) for sophisticated threats like state-level actors. I hope LLMs continue to be unable to do so for the near future (I just tested and LLMs can't do it with a simple prompt).

Which is why I think privacy safeguards still work quite well even while being technically mostly bypassable.

So probably tools like this should be kept private if possible.

Looks like the site was taken down. I was curious who wrote the most like me, given that I don't have any alt accounts, but I guess weighing privacy over my curiosity is a good thing in the big picture
Entropy is unfortunately a very bad metric to estimate if these identification techniques will scale. Plugging my work here as example: https://www.nature.com/articles/s41467-024-55296-6
"Pseudpocalypse"... why not "pseudocalypse" or "pseudoapocalypse"?

As an aside, it's always surprising to see how English speakers split Greek words like "Apocalypse". That is to say, they always split them in the middle of Greek syllables, or just drop letters like "pseud[o][a]pocalypse" and often in a way that ends up sounding clunky and weird even in English.

Can't think of other examples now. Brain going to sleepzzzz....

Edit: oh wow there's actually a word for that:

https://en.wikipedia.org/wiki/Libfix

OK, now I go to sleep.

"pseud" is it's own atom here, see definition 2 in https://www.merriam-webster.com/dictionary/pseud

It's not trying to say "something like an apocalypse" it's saying "an apocalypse of pseuds"

Let's see, expanded it's "the (or an) apocalypse of pseudonyms." In Greek, this would roughly be "ἀποκάλυψις ὧν ψευδωνῠ́μων." It's not impossible for a genitive like this to become a word, possibly "ἀποκάλυψευδώνῠμοι" (think of "φιλέλλην") -- although it's a little fanciful. There may even be precedent for a construct that puts "ψευδώνῠμοι" first. But ultimately I agree that the word the article offers seems tortured.
English speakers even mix latin and greek roots!

From a t-shirt I saw once: "I am against polyamory! It should either be multiamory or polyphilia!"

It's a speech and attention thing. With the American accent you weigh the pseud part higher than the -onym because it's the downward part of the word. pseudocalypse highlights the -ocalypse on the upswing which detracts from the pseudo- so to maximize traction use both recognition highlights pseud- and -pocalypse.
To achieve this level of total surveillance it would trillions and trillions of dollars to be invested in a massive roll out of data centers. I just dont see it happening.
Right now… it would have been science fiction 10 years ago and not a trillion dollar problem.
> I just dont see it happening.

It is though.

Perhaps remarkably, the US intelligence community is funding work on this problem in the open, as part of IARPA HIATUS [1].

The program's initial phase is winding down now, so some of the performers' papers ought to be hitting the ArXiV before too many more months.

[1] https://www.iarpa.gov/research-programs/hiatus

Related articles:

- Claude knows who you are: https://www.lesswrong.com/posts/Jkb4CBB7rf4XYP5eb/claude-kno...

- ~ Opus 4.7 is the first model to correctly guess who I am based on unpublished articles: https://x.com/KelseyTuoc/status/2044962428547695007

I wondered something almost the opposite the other day. With consistent interaction with chat bots and AI output, their linguistic ticks are surely bleeding into every day speech ("smoking gun" and other turns of phrase). What if our linguistic output starts to conform?

> Do they, incorrectly, position their adverbial clauses? Underrated line.

I'd seen the articles making claims about LLM-idiom being unconsciously imitated by humans, and reassured myself that _I_ couldn't possible be so impressionable.

Then during a presentation I was making today, a coworker wrote in chat "If <scubbo> says 'You're absolutely right' one more time, I'm checking his house to see if we're actually talking to an agent".

They got me.

Our linguistic output definitely IS conforming. It would take our individual sustained efforts not to, and even then it works be hard for is to not to… language is a meme and AI is very good at infecting us with all kinds of memes
That whole paragraph is clever :)
I've gotten several dozen accounts banned on reddit over the years. I don't remember the names of them all. It would be hilarious to me if someone mined the archives and stitched them all together.
As long as you're fine with losing your distinctive voice (which should be taken as table stakes for people who value anonymity to the extent of worrying about this), it's perfectly reasonable to use tools to stymie stylometry. I'm not even suggesting using an LLM (which may or may not be sufficient, it would be difficult to verify either way); it would suffice to have a tool that rewrites your prose (or analyzes it and flags it for manual correction, etc.) if it isn't written in, say, a form where every sentence is primitive subject-verb-object, limited to the 1,000 most common English words, with no contractions, idioms, or exotic punctuation. Yes, this doesn't completely eliminate all possibly identifying bits of entropy, but it would more than suffice for hiding in a crowd ("hiding" in the sense of obviously standing out as someone trying not to be noticed, and as long as you're also careful about your opsec in other ways, like time of posting, etc).
So who is Satoshi Nakamoto?
then you'd imagine the inverse is possible: fabricate text that matches a fingerprint - poisoning someone's (psuedo-anonymous) reputation.
I know I'm a weirdo, but I've been posting under my own name since the 90s.

I guess I always figured that nobody gives a shit who I am.

There's plenty of other people on the internet with identities worth stealing. You'll probably be fine. Still, it's not a lottery I want to play.
those deanonymization vectors (stylometry and input device biometrics) have been known for a while in the privacy community. re: stylometry, AI rewriting will be useful once it becomes fast & cheap. but research shows that trying to emulate another writing style is already effective. for input device biometrics, there is the kloak algorithm. and which ships with qubesos https://github.com/QubesOS/qubes-gui-daemon/pull/149
I wonder if this could be used to unmask satoshi. I remember a piece about applying stylometry to satoshi's writing, but they just compared him to the usual list of suspects (finney, back, etc.)
One would think that it could be possible to make a tool that takes some text and "anonymizes" it by making it a little more standard and boring (uniforming punctuation and sentence structure, changing words with some synonyms, etc). Maybe wouldn't make it particularly compelling, but would be valuable for political dissidents and other people with a high threat model.

Does anyone have some tools to share?

I wonder how this would work if, for example, someone who never uses AI for writing decides to sometimes use AI for a sentence or string of sentences here and there. Or vice versa, an AI addict makes themselves hand write half the sentences sometimes.

Basically, putting a pebble in one's shoe to fool gait recognition, what's the equivalent thing for defeating stylometry?

That gives a new reason that people should read LLM generated articles. For privacy purposes, if writing is run through Claude to add LLM-isms, then it's harder to deanonymize the writer. If Edward Snowden were posting files online, he'd want to post his writing that way too.
So anonymity of written speech is toast. We should, however, strive to preserve other forms of anonymity. For example, donations given to political causes should be kept confidential. Let protesters wear masks up to the point where they break the law.
There's a new 3Blue1Brown video hinting in the same direction.

https://www.youtube.com/watch?v=GlYgs6v2YfU

Sooo, this implies such deep profiling hasn't been in use for a decade for target advertisements. 500M is not that large a number with the amount of traces we leave behind online :)
What a phenomenal article, wow.

This is one of the best things I've read on this site.

I have a side project/experiment that's tangential to this (wafertown.com), so my interest is 2x the usual.

so really, all these people using LLMs to comment aren't being lazy! No, they're using cryptographically linguistic security to ensure untraceability!
yes but no, can't we just ask AI to sufficiently shuffle our words or for algos to do so?

"boom", pseudoanonymity (spell?) restored?

See the classic Gwern post: https://gwern.net/death-note-anonymity

Which quotes Tao on using deliberate disinformation to preserve anonymity.

> …one additional way to gain more anonymity is through deliberate disinformation. For instance, suppose that one reveals 100 independent bits of information about oneself. Ordinarily, this would cost 100 bits of anonymity (assuming that each bit was a priori equally likely to be true or false), by cutting the number of possibilities down by a factor of 2100; but if 5 of these 100 bits (chosen randomly and not revealed in advance) are deliberately falsified, then the number of possibilities increases again by a factor of (100 choose 5) ~ 226, recovering about 26 bits of anonymity.

Intentionally adding writing "tics", scheduling posts to appear between 2am and 6am in your timezone, or pretending to have a different gender/location/age should help a lot in staying pseudonymous for a while longer.

This always seems theoretical. Has it happened?
> A stronger conjecture is that we’re heading towards a sort of generalized pseudpocalypse. Perhaps, in the future, if you interact with the world through essentially any high-bandwidth channel, then you identify yourself. Say you wear a mask in public and only speak by sub-vocalizing into a voice changer. That’s fine, you’ll still be identified using your body shape, gait, or chemical signature. Or say you don’t like your car being tracked everywhere, so you stop carrying a phone and you somehow convince lawmakers to ban license plates. No problem, your car will still be tracked using tiny scratches or unique pinging sounds from the engine. Or say you don’t like being tracked on the internet, so you lock down your browser profile, buy stuff only with Monero, and connect through a chain of three VPNs. That’s OK. You’ll still be identified through how you wiggle your finger as you scroll down the page. We’re all just too unique, and the information theoretic limit is coming for us.

Forensic research, NSA, Palantir…

Btw 42. Sleep, eat, have sex, have fun, be useful.

Huh, I just realized he (she?) was pseudonymous and not Matt Might[0] this whole time. Oops.

Also, I wonder about this analysis in the age of AI slop. I wonder how much that removes the identifying bits, vs how much carries through of the original prompt (e.g. topic and guidance). It's interesting that a pseudonymous blogs might take on very generic Claude-voice, which could be worthwhile if the topics were interesting, but could also just be a completely humanless bot.

[0] https://matt.might.net

Eh isn't authorship linking a whole field of study? Yes it is, called stylometry. Here's a review paper form 2006 [0]. Its an old subject, with literally thousands of papers. Really wished the author had taken a cursory glance at the literature.

[0] https://link.springer.com/chapter/10.1007/11889342_20

a lot of these comments have no understanding of basic statistics. for example, did you know, at just 30 words of grammatically acceptable text, at least in english, youre already at more possible combinations of phrasings than atoms in the universe. you cant just throw a problem like that at LLMs and expect them to "just work"

folks should also look at burrows delta - i forgot which books but some folks were able to identify a ghost writer by stylometry alone.

ive said this on many threads, you cant just have text output from an llm (regardless of style / "pseudonym") and have it be "unique" because the nature of the transformer model itself is literally present in the output words. it will be detected as llm output every time. it has to be!

for pure anonymity, i suppose then it is an answer... for the actual craft and art style of writing it is not

The best mitigation a person has is to have an LLM 'flood the zone' with slop based on her writing style.

So she has one comment on the internet admitting she cheated on her taxes, another copping to an axe murder, another revealing she's the one and only D B Cooper.

One potential solution here I suppose is to make deanonymization or contributing to it a serious crime. I doubt that will happen in most places, though, since it’s often the government that wants to do this to its own citizens.