https://arxiv.org/abs/1503.02531
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
https://arxiv.org/abs/1503.02531
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?
> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.
It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
... but Chinese SOTA foundries directly using distillation as fair game.
I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
What is more reasonable:
- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.
- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.
Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.
You can't distill what you are not given - simple as that.
Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.
I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.
Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.
Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.
It's hard to draw the line.
But the Chinese models are absolutely distilling - and would not be competitive without this distillation.
At the same time, there's a lot of real innovation and regular building going on at the same time over there.
I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.
BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.
This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term.
Moreover - the 'reasoning traces' are not required for distillation at all.
Finally - it's entirely possible for them to have used Fable for later stage fine tuning.
It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'.
This is just straight-up factually false.
The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up.
Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude".
Don't make stuff up to suit a political agenda. It's extremely dishonest.
Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable.
Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for distillation purposes, it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.
> specifically designed to be useless for distillation purposes
No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.
> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
I did not claim that. Read my comment again:
> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
Because apparently I have to spell it out:
The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.
I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?
I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic.
You're gaslighting me. I did nothing special at all, and there's ample evidence of this happening to others on Twitter.
> you don't need to be paranoid and assume they must be getting it all direct from Anthropic.
Nowhere did I say that. Stop lying about my words.
To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source.
It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source.
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs.
> To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it?
there's no creative work between the weights and the tokens being made.
whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images
You do realize the 'investors' are the one's who 'own' companies and therefore the IP?
One person can a corporation be.
> ... but Chinese SOTA foundries directly using distillation as fair game.
As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.
LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..
If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.
If not there is not there is no grey zone whatsoever.
for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.
i don't use llms for that reason.
> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.
> ... but Chinese SOTA foundries directly using distillation as fair game.
two wrongs don't make a right, but the irony is at least something.
> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
the corpos can get fucked as far as i'm concerned.
> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
*only in the US.
I'm just nothing that HN rhetoric is contradictory.
But this:
"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.
If it were some true injustice, then fine, but that is clearly not the case.
There is ample room to contemplate that even copyrighted works could be considers fair use as training material.
"the corpos can get fucked as far as i'm concerned."
Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.
I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.
Periodic reminder that HN is not a collective or a singular entity and is actually a bunch of different people with different opinions. Often the people with the loudest opinions get upvoted to the top - and often the "side" represented at the top is different from thread to thread.
There is no second “wrong” here.
Model outputs are not copyrightable. I think that was already established?
Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic?
If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business with them. Maybe even sue them in civil court for breach of contract (again nothing to do with copyright infringement though)
anthropic are essentially saying in this tweet they believe a moral wrong has been committed against them -- "unacceptable behaviour" etc.
plenty of people have been vocal about the fact anthropic have committed moral wrongs at scale in building the products in the first place, with the question of legal wrongs still being worked out.
so, two moral wrongs. legally, fuck knows.
I mean you are right in a way of course, it’s just a matter of degree and perspective, though. If one thing is moderately morally wrong and the other is potentially lightly morally wrong I don’t think it’s fair to equate them.
To me the situation is a bit like Google coming out and saying that its morally wrong for someone to build a competing open operating system on top of Android while stripping all Google services and “stealing” their ad revenue. Just seems silly and hypocritical.
And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.
No wonder people are fed up with this behavior.