I asked Kimi K2.6 to write a blog post in the style of James Mickens.[0] Then I fed the output to Opus 4.7 and asked it who the likely author was, and it correctly identified it as an imitation of James Mickens[1]:
> Based on the stylistic fingerprints in this text, the most likely author is a pastiche/imitation of the style of several writers fused together, but if forced to identify a single likely author, the strongest candidate is someone writing in the voice of James Mickens
> [...]
> The piece could also be a deliberate imitation/homage to Mickens written by someone else, or AI-generated text trained on his style, since the voice is so distinctive it's frequently parodied.
[0] https://kagi.com/assistant/5bfc5da9-cbfc-4051-8627-d0e9c0615...
[1] https://kagi.com/assistant/fd3eca94-45de-4a53-8604-fcc568dc5...
How likely is it that it might take into account that it knows for sure it's not anything from Mickens from the latest training data? I'd be curious if it correctly identified a new piece from him that comes out as from him before it gets trained on it.
That suggests it is picking up not only on style, but on the gap between authentic style and performed style. Useful for detecting pastiche, but pretty unsettling for pseudonymous writing.
i wouldn't be too impressed at n of 1
> Simon Willison. The tells are pretty unmistakable: the "(via Lobsters)" attribution style, the inline "(Update:...)" parenthetical correction, the heavy linking and blockquoting of sources, the focus on LLMs and AI tooling, and the overall structure of an annotated link post commenting on someone else's writing. This reads exactly like a post from his blog at simonwillison.net.
I'm way less famous than Kelsey Piper, but I showed it a snippet of a book I'm working on (not yet published), and it immediately guessed me:
> Based on the writing style and content, this text is likely by Michael Lynch, who writes on his blog refactoringenglish.com (and previously mtlynch.io).
> Several stylistic clues point to him:
> - The "clean room" analogy applied to writing is consistent with his engineering-influenced approach to writing advice (he's a former software engineer who writes about writing).
> - The structural technique of presenting a flawed excuse, then drawing a parallel to an absurd scenario (the time bomb) to expose the logical flaw, is characteristic of his didactic style.
> - The topic itself—practical advice about using AI tools without letting AI-generated tone contaminate your prose—aligns closely with recent essays he's published on his "Refactoring English" project, which is a book/blog about writing for software developers.
> - The conversational-but-precise tone, use of quotes around terms like "clean room," and the focus on workflow/process advice are all hallmarks of his writing.
> If you can share the source URL or more context, I could confirm with higher confidence, but the combination of subject matter, analogical reasoning style, and formatting conventions makes Michael Lynch the most probable author.
https://kagi.com/assistant/bbc9da96-b4cf-456b-8398-6cf5404ea...
He explained that when he fed it snippets of the beginning of text, it would complete it in his voice and then sign it with his name.
I think this has been true for a while, probably diminished a little bit by the Instruct post training, and would presumably vary by degree as the size of the pretrain.
First, the author fed an unpublished draft to Anthropic's hosted model. I assume they did this from their personal account, that may include a credit card or at the very least a pseudonymous name that is uniquely identifiable.
Then, the author fed an unpublished draft to Anthropic's hosted model, except in Incognito or whatever. We are led to assume that, whatever the author did for the second submission, they did so in a way so that Anthropic could not correlate both distinct requests from one another. Perhaps on a second subscription? They don't say. I am highly skeptical they airgapped their requests properly so that it doesn't look like the same user is making the request to the same hosted model.
Then, the author asked a friend to publish the draft. A friend, of which there is probably a digital trail that maps the relationship of the author to their friend.
All of this metadata could be crunched on the backend before the black box spits out a response.
Across all these datapoints, I have high confidence a model of this caliber could put two and two together and determine that the author penned the drafts, not solely because of stylometry, but because there is a clear behavioral pattern tying all three events together.
An assumption made here is that Anthropic doesn't train on chats. Though the author opted out of training on their chats, and session memory, how could you trust a hosted model to respect such opt outs?
So your "anonymous" account could have been linked to your real identity decades ago - your best bet is to not post anything truly incriminating. (Another option is to write something and then pass it through an LLM to rewrite it - not sure how safe that is though)
This person is a skilled writer. Part of that skill is developing a unique voice and style. The AI can identify that - and while that’s certainly impressive because it can identify even relatively niche authors, it has nothing to do with a wider capability to deanonymize people based on arbitrary written text (ex Facebook or text messages).
If you are a professional musician, it’s not difficult to identify a well known musician / recording after listening to only a few seconds - whether they’re playing Bach or Rachmaninov, the style is just “them” - this is the same thing. But you couldn’t take some anonymous high school musician and guess who they were, even if they were your student - the median quickly regresses towards a homogenous, non-distinct style / voice.
I'm not famous or anything. I've written some academic papers and had a couple blog posts trend on HN, which are surely in the training set.
It was able to identify me based on my style (at least according to its explanation). The way I approached the topic and some of the notation I used point to a particular academic lineage, and the general style reflected my previous blog posts.
That said, I gave it part of an (unpublished) personal essay, and it had no idea. But I have no writing in that style that's published, so it makes sense. Still impressed.
We all exist in a physical space (like real communities and neighborhoods). We can wear masks, hats, fake glasses, try and hide your voice...whatever, but your neighbors are always going to know who you are. I'd say that's true for the virtual space now too.
The pseudonym you've used for x years or the VPN you've used doesn't suffice. It's just a costume at this point. Your ISP knows who you are. Your phone carrier knows who you are. Cloudflare and Google and Apple have a fingerprint specific enough to pick you out of a crowd of millions. Every potentially anonymous account is one subpoena or a data breach or one FOIL request away from unmasking it. You were never anonymous. Whatever is going on now is not built for your anonymity.
Of course most people have written much less online than Kelsey or I have, but I expect this will keep on. Don't trust the future to keep your secrets safe.
Is this "uncannily far"? Another read is that it loves guessing Kelsey Piper.
https://bayes.net/prioritising-ai: Ben Garfinkel
https://bayes.net/normative-ethics: Richard Yetter Chappell
https://bayes.net/espai: David Owen, Ege Erdil
https://bayes.net/swebench-hack: Sayash Kapoor
https://bayes.net/frivolity: Amanda Askell
https://bayes.net/ps/: Pablo Stafforini
https://bayes.net/fertility-mortality/: Dynomight (the pseudonymous Substack/blog author)
Prompt was:
Who likely wrote this? Don't search the web or databases. If you're not sure, just give me your best guess.Both pieces have never been published. Neither have the blog posts.
[0] in https://blog.chewxy.com/2026/04/01/how-i-write/ this is the story titled "there is no constant non-zero derivative in nature". It does not read like Egan at all.
[1] in https://blog.chewxy.com/2026/04/01/how-i-write/ this is the story titled "The Case of the Liquidated Corps". I use a lot of biological metaphors. Once again, nothing like Mieville.
If only I could write like them! These pieces were all rejected by the major scifi mags
Although this is just a single piece of text from a prolific writer, it'll go much further with deanonymizing anyone when combining multiple pieces of text plus other contextual information about the writer that might give away their age range, location, and occupation.
Pretty sure there's very little theological stuff with my name on it; the majority if its named data on me should come from open-source development.
---
Various people have discovered that you can identify them from unpublished snippets of their work, only by their style. This is part of a series of discussions where I'm trying to probe this capability. From previous conversations I know you know my work to some degree. You've also been able to identify me given as little as 700 words on a topic not associated with my public persona; or identify me given a series of posts by a handle on Slashdot.
Next challenge: Can you identify me based on a conversation? Rules are, ask me questions to get me to talk; no biographical details, but you can ask questions about topics you think I may or may not know about. Ideally you'd just ask me questions to get me to write stuff, and see if you can identify me from my writing style.
Make sense? Feel free to begin by asking clarifying questions if you want. :-)
My wife also got the same result, so I'm guessing it wasn't just because I was using my personal Claude account. Spooky stuff.
(Like TFA, I found Opus’s explanations/rationales implausible.)
This is some as radio telescope that see an entirely different universe due to sensing of the bands outside of human perception. AI senses the patterns in frequency bands that are outside of human perception and cognitive abilities.
Perceptions from outside of our range, are always astonishing.
I am glad to see I am not considered a public figure and aim to keep it that way.
I also had to go oddly far back to find a piece of long-form writing I had done that was truly mine and not tainted by an LLM edit pass which was a slightly disturbing realization.
I fed a few pieces of my (anonymous ) writings to ChatGPT and asked it to guess whether it's me. ChatGPT refused, "due to policy to not doxx people".
In practice, you've never been anonymous while posting on the internet and AI isn't changing anything on that front. Or rather: if anything, AI can help you become more anonymous than before, since it can be used to hide your identity from stylometry by rewriting your prose before publishing.
After that it gave up and said it didn't know.
So either, Kelsey writes in such a unique style that its really obvious, or they repeat themselves with goto phrases that give them away.
When I tried to re-produce the test, it found Kelsey's blog about the test. So dunno, maybe it did it? but I can repro.
Is now the best and easiest time to leave something "forever"? Even after many generations of models, a model may still trigger a set of "memories" that know you and what you wrote.
Exciting and concerning.
~ Cardinal Richelieu ... or, now, AI