back

by malshe·7d ago·view on hn ↗
This is rich considering LLMs can't even write like an average person let alone famous authors
5 comments
LLMs cannot write dialogue between two characters where one person knows a fact and the other person doesn't. It will consistently make the character who shouldn't know something bring up said fact and I have not found a model which will not do this. All these years of LLM research and it fails grade schooler level logic, if it is even slightly abstracted behind a story.
A good plenty of counterexamples on eqbench.com.
this is rich coming from the company that shreds books for training data
There is a really weird perception of the importance of books as if there is a important property to be preserved in every book.

This was hilighted to me by a librarian friend if mine, She said they destroy books all the time in her job, and got rather angry when people suggested that it was intrinsicly bad, because there is nothing sacred about simply being a book. It is about the replacibility. They had a program where, if they destroyed a book they would receive a replacement from the publisher for a tiny cost. If a book got damaged it would be much cheaper to get a freshly minted replacement than it would be to spend time repairing.

If there was a glitch in the matrix and 30 billion Gutenberg bibles suddenly fell in the middle of the Amazon(the wet one), most would be destroyed as an environmental hazard.

I see a lot of comparisons to book burning when the AI scanning is mentioned, but the point of book burning is a symbolic act to indicate that the information within the book should not be shared.

Destroying the book to capture the information within sits at the polar opposite reason for destroying a book.

The difference here I think is that the company (Anthropic, not OpenAI) was specifically selecting rare books for training (and subsequently destroying). I think the value proposition is a much different one in this case. In your example the books are still available in high enough amounts that replacing it is cheaper. If the book is rare enough to not have been digitized for training data yet (which is why Anthropic is interested in it), keeping that book around is much more important.

If Anthropic subsequently released the digitized version of it we’d be having a different discussion, but for obvious reasons they’re not.

I have heard allegations that they might be destroying non replaceable books, but only as a hypothetical.

Replacibility is the standard tbat matters in this which I mentioned in my post. I haven't heard of a specific instance of an occurance of destruction of an irreplaceable book, but I'm willing to look into cases if you have a link.

> I have heard allegations that they might be destroying non replaceable books, but only as a hypothetical.

> Replacibility is the standard tbat matters in this which I mentioned in my post. I haven't heard of a specific instance of an occurance of destruction of an irreplaceable book, but I'm willing to look into cases if you have a link.

I think you can assume they're destroying rare or even irreplaceable books, unless they can show they've implemented careful processes to avoid doing that.

But given they're all basically SV startups, it's very unlikely they're doing anything except the minimum effort to get what they want now.

For me it’s just as unbelievable that the second hand book stores would sell rare or irreplaceable books at bulk prices. If they do they are really not good at what they are doing and I wouldn’t blame any startup big or small to destroy it, especially after preserving it digitally.
> If they do they are really not good at what they are doing and I wouldn’t blame any startup big or small to destroy it, especially after preserving it digitally.

1. That's fucking bizarre logic. "It's ok to do bad thing if someone didn't stop you."

2. These startups aren't "preserving" it digitally. The destroyed physical copies would have likely outlived whatever datasets they're creating.

Rare != valuable.

They are "betting" pennies in their books.

Rarity equating monetary value is a common logical fallacy often found with stamps and other collectable cultural artifacts (details would disgress too much here, see Pokemon cards or whatever).

And rarity vs quality ir utility is even harder.

But value is 1) not determined by scarcity and 2) neither by "unique information content" (because that's by definition not measurable without a panopticon) and 3) nor by "creative value" per se 4) it is indeterminate whether some detail in a work might prove useful later, in general. Even for scientific papers.

They just can bet on all fronts and for training material it's 100% rational to do it.

Apart from that, to briefly disgress into the discussion of "cultural value": many cultural works we hold as very important today were sold for cheap, if at all, when the author died.

Music scores are good examples.

But back to the broader "value"/rarity angle:

The historical value of some odd self-help or DIY instructions book from the 80s is hard to determine! It might happen to be able to provide a secondary source for something, or even act as primary source for some other fact of the time that is tangential to its subject.

None of the qualities that make trivial literature interesting to, for example, researchers, must correlate with market value (at the time of writing or at the time of research?).

Sure, there are truckloads of "unique" books from even 200 years ago that are neither culturally valuable nor sellable right now.

But these things can change quickly.

Imagine you want research details about old retail products.

Or the usual street price of something in some city at a specific time.

Or the inhabitant of a certain rentable space at a certain time.

Or some specific detail about replacing the parts of a certain type of car that's no longer around (is this valuable then?).

Humans create endless complexities, and monetary or evolutionary value is not coupled anymore to step-changes such as "how to light a fire using X,Y,Z".

We're deep down the rabbit-holes of our own ephemeral shaping of our world.

What word was commonly used to describe a sandwich in 1925 in Hungary.

Or what problems did people commonly encounter when repairing car doors in the 1940s in England?

Whatever. I can't make up a good example right now, but market value does not imply quality; and even quality (according to some cultural norm) is not universally measurable.

They did it due to some weird copyright argument about copying and destroying the original seeming like a transfer and not copyright violation, which is insane and apparently is mot always interpreted in said way. So they are literally taking a copyright gamble they might lose.
It is fair use to train on books you buy. The anthropic case established that. The destruction of books is just cutting the spine so they have individual pages to scan.
Destroying a book to argue fair use is not accepted as a fair use act is what I was saying, so they're needlessly destroying rare books.
They are not destroying any books to make any claim. They are doing it so the pages fit in the scanner.

Without an example of an actual rare book that is shown to have been destroyed, all there is is the is the claim they have the capability to do that.

The closest I have heard is the possibility of rare books purchased in bulk orders of cheap second hand books. This is a similar risk to melting down someone's wedding ring from a batch of scrap metal. It could theoretically happen, but the knowledge of the items presence is absent and it was not requested.

What other allegations have there been? I'm prepared to look at actual examples of you have them.

I think the important part of books is the knowledge in them. Not the binding or anything else physical. With some exceptions of course and sure there are things to be learned about physical attributes that don't translate to the digital version, but 95% is in the content.

> If Anthropic subsequently released the digitized version of it we’d be having a different discussion, but for obvious reasons they’re not.

I don't think you can scan books and just release a pdf. Someone correct me if I'm wrong. Even if you can, I'm sure it's a gray legal area.

99.99% of printed books are worthless.
99.99% of stuff is worthless for most people.
Also what's worthless to you is very much not worthless to others.

A pretty obvious example is genealogical archival records. My family only cares about a few dozen pages out of millions for each US census, but every family cares about a different few dozen pages.

I agree for famous authors. But I also believe you are overestimating the average person when it comes to their ability to articulate their thoughts.
Just another propaganda lie, implying that they could imitate any style.
The early models could to an extent. They removed it, probably with RL, because the feature reveals the inherent plagiarism.

That is also why ChatGPT blocks it now. The plagiarism is still there of course, just hidden.

They can imitate distinctive authors.
Even LLMs from years ago could accurately mimic the style of any sufficiently famous author.
LLMs from years ago weren't aligned to death like these days. It's hard to get rid of "AI smell" than years ago.
Yeah and it's degraded significantly since then. Older llms were still mostly language focused and had a lot of latent knowledge about things like writing styles. Now it's crowded out in favor of agenetic work, programming etc.
No, the latent knowledge is larger than ever, as verified by many benchmarks (and instances like people being shocked by truesight of obscure forum posters). This is 100% a chatbot personality/alignment/post-training thing.
> truesight of obscure forum posters

What does this mean?

Truesight is a spell in D&D that allows you to see the true being behind an illusion or other deception.
It's almost like a claudeism
Gwern, who posted that comment, is one of the early popularisers of the term. Janus used it in at least 2023, and Gwern in at least early 2024.

If anything, you probably have the causation backwards: Claude talks like a LessWrong poster.

I’m surprised no one has distilled a 4o model yet that’s really good at prose. No money from enterprise users I suppose.
At best it can use some of the same vocabulary. Often not even then.
Can anyone point me in the right direction on how to get these models back? I really like the older ones for NPC dialogue for TTRPG campaigns and the like.

I just looked into it and I think I can still find GPT 3.5, but wonder if it too has been trained out of usefulness.

The characters a styles were quirky and not always perfect, but added a lot more flavor and character nuance that could be corrected with editing.

I have no experience with style imitation, so perhaps I'm missing something here. I asked Kimi K3 to rewrite your comment in the style of Ernest Hemingway:

> I would like to get the old models back. If anyone can show me the way, I would be grateful. I liked them for NPC dialogue, for TTRPG campaigns and the like.

> I looked into it. GPT 3.5 is still around, I think. But perhaps they have trained it out of usefulness. It is difficult to say.

> They were quirky, the old ones, and not always perfect. But they had flavor, and their characters had nuance, and what they got wrong you could fix in the editing.

Full conversation with thinking: https://paste.ononoki.org/?2fc049fe1ae7302b#GjzuEn8VmaYJ565M... Password: kimitest

This is really neat, but if you'll notice, it just kept the same words and rewrote it a few different ways. I guess what I'm actually thinking of it's when the model temperatures weren't quite dialed in and you would get some wonky stuff that was genuinely useful for character exposition and giving a particular character a "voice".
How token efficient is kimi? How much do you average each month for example?
You can pay Microsoft they were exclusive host of GPT models.
The base models were really good at this. These days even the "base" models are full of awful synthetic data.