back
910 comments
The 2 things people need to remember:

1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.

2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.

https://thereallo.dev/blog/claude-code-prompt-steganography

Why should I trust a US company more than a Chinese one?

1) I'm not using AI to bicker over fringe political shibboleths.

2) I don't think it is any more or less safe to put my code on a Chinese server versus an American one. A Chinese provider also isn't liable to spy on me for the feds, as OpenAI and Anthropic certainly do.

Regarding point 2, I don't trust my data being safe running inference on model creators api, but neither do I trust US providers. Both use it for their own benefit, the only difference is the country of origin. The US has a lot more legal safeguards for this but I don't trust they don't do it regardless.
It’s perfectly fine for the “West” to influence the world though, right? Or is it only a problem because… they’re Chinese?
On (1), models are an aggregation of large volumes of data sources, whatever the culture producing them, you'll get the average bias of that culture.

We see that on what minorities are associated with inside the model, or how things that aren't online will have a completely different weight. Or how 2/4/5/8ch or X will be disproportionately present in specific models despite being the places where facts go to die.

But hasn’t it been the same with US since the WWII? US has influenced the world through various ways sometimes even weaponising human rights to spread American superiority and its narrative?
but you are not using chinese models to learn about taiwan, you use chinese models to do everything BUT learning about taiwan, so no issue there
> China can (and does) use the models to influence the west

OK now that's false information.

You can uncensor, tweak or fine-tune open-weight models, but not so easy on a proprietary model from some cloud provider.

1. Get the free Chinese model. 2. Jailbreak it 3. ??? 4. Profit?
The post hits it spot on with unequal access to the models in terms of security. I'm developing OSS where security is important for the user ... but the frontier models like GPT 5.6 and Fable flake out and state that I cannot get the info/access.

This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly.

For OSS, this is one of the most counterintuitive experiences I have ever had. More than ever I'm convinced that open weight and open pipelines models are 100% critical for progress on the AI and societal fronts.

People seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."
I am afraid — if Chinese models go mainstream it has a clear way of pushing its narrative way beyond its otherwise borders. More like a Trojan horse it is for the Chinese.

Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question

https://x.com/jinen83/status/2079406993979383902?s=46&t=D7hQ...

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation

Sounds great to me; live by the sword, die by the sword.

The article makes a point about agent harnesses being sticky (the supposed moat). I have been building my own agent harness for a while, and I can tell with confidence that the harness almost does not matter, the entirety of the AI magic is the model itself. The harness can be almost barebones (like, for example, mini-swe-agent used for benchmarks), and yet the model still does the task just fine.

So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).

The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models for free. If the frontier labs are forced to cut prices and join the race to the bottom in token prices, these valuations are unjustified, and VCs will face enormous (paper) losses.
Lets do "who's afraid of US models" version:

* Me, as an individual, because I might not be able to pay price hikes, because my revenue (salary) is much lower than what they want and I can't support my expenses via huge bank loans.

* Again, me as a new entrant to the industry, LLMs are basically pay-to-play games, again related to price hikes, new entrants might not be able to afford paying those prices 24/7 - which you need when learning new things.

* Any non-US company, US can block the models which can disrupt the whole business.

* Even some US companies, for example if you operate in EU and EU somewhat changes their mind and follow the ICC and require you to stop working with Netanyahu (war criminal as per ICC), then following laws in EU, might create trouble to your whole business.

> The defining characteristic of a commodity is that it is fungible: a gallon of oil is a gallon of oil; a ton of copper is a ton of copper; a bushel of wheat is a bushel of wheat.

The concept of “commodity” as defined above is a model, a simplified abstract representation of reality, but that does not match the reality perfectly (the map != the territory).

The author claims that a token isn't literally an ideal commodity, but neither is oil or wheat, many factors influence their real value (intrinsic properties, location, available storage at production, expected delivery date, etc.) so that no two gallons of oil in different contracts have the same price.

Is treating “tokens” as a commodity a worse model than treating oil this way? It depends who you ask! I'm pretty sure that a chemist working at a refinery would be more happy to see tokens being felt with like a commodity by his company than if they started viewing crude oil like one.

(Overall, there's way too much economism in that post, and way too few facts, and as a result the argument makes very little sense, the author basically wrote that both OpenAI and Anthropic are drowning in cash right now because compute scarcity means the price must be significantly higher than the marginal cost…)

Releasing open weights that can approach frontier level intelligence (irrespective of number of tokens burned) is just a way of telling the world that anyone, even China, can serve frontier level inference if they have the chips and warm shells to do so.

What is stopping China from gaining a majority market share, then, in terms of serving inference?

AI Sovereignty -- yes

Cybersecurity concerns -- yes

Latency -- no, unlike previous emerging IT workload types , inference does not have strong latency requirements. eg 1s of additional network latency doesn't matter to a 15 min, 10-turn agent session.

Cost -- ultimately this comes down to a nations ability to plug chips into warm shells. which forks into geopolitical / trade on the chips side and energy scalability and modularity on the warm-shell side. Even if you call geopolitical / trade a toss-up, China has the US beat HANDILY on the energy front, yearly they are deploying 10x power to their grid relative to the US, which is shooting itself in the foot at every possible moment.

IMHO chip tech will travel across borders, absent a breakthrough in analog inference, energy scalability will ultimately dominate.

I operate an analytics site (pretty big one B2B where client's backend feeds data into our system), and we see tons of traffic originating from northwestern China (Xinjiang) from Shenzhen Tencent Computer Systems Company Limited.

There are also half a dozen other companies from China continuously hammering our clients’ websites.

I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.

Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region

credit:

'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopoc

Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.

Assuming that China only distills is a huge mistake.

It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.

Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.

The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.

Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.

This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....

There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.

> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users.

My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.

[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]

“ Let the frontier labs win by being better; don’t let them define safety or security, or pull up the ladder of humanity’s collective knowledge”

Love this.

I'm rather scared of US models - if Anthropic was the only AI provider in the world, it's easy to see that common people would have no access at all. Thankfully there is OpenAI which compete 1:1 with Anthropic (at a slightly lower cost) but most importantly the Chinese models keep Anthropic, but also OpenAI in checks.

And I'm saying this as someone working for American companies.

Try different harnesses people! I am actually preferring Chinese models at a fraction of the frontier price for coding. Yeah you need more tokens per unit of work done, but it is way cheaper still. Using CC/Opus as a staff eng / frac CTO. And Hermes/Chinese model as hopefully my team of mid levels. This way I can make good use of pro plan and then get cheap Chinese tokens for the rest and not hit a RL and know it can scale up. plus choosing your model is so cool and some are a lot less verbose.

Hermes is a better coding tool IMO. I can't put my finger on why but it just feels better. Maybe being true yolo helps.

I think in general rest of the world needs to take notice (not saying afraid), starting with the US. It cannot be taken for granted that China's frontier labs will be a few months behind. They might be at par or exceed.

The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.

At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.

Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.

Excellent article; the argument towards the end for allowing distillation for US companies is compelling:

> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?

> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.

The U.S. "executive" class is so obsessed with the "exploit" part of the explore/exploit cycle that it's very clear they are prematurely closing advancement. Better a little money and power for them now than a lot of money and power for their country/humanity.

This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.

You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.

According to openAI's own @deanwball: Even OpenAI isn't buying this distillation talk:

https://xcancel.com/deanwball/status/2078133895766114412#m

> This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?

This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.

And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".

One thing I have not seen mentioned between Chinese AI vs US, population.

China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.

Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.

So I believe, China will end up owing AI.

People who claim that the Chinese open weight models have some type of manifest advantage don't realize that the close weight models have a huge advantage as well: the researchers from OpenAI, Anthropic, Google, xAI, Meta are not dumb, they can read the white papers written by DeepSeek, Moonshot, etc, and they can inspect all those architectures and they can pick and choose the best tricks there are out there, and of course, they have access to their own in-house secret sauces.

Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.

But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:

1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.

2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.

3. The frontier labs are also investing more and more in building an ecosystem around their models.

I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.

[1] https://www.anthropic.com/careers/jobs

"distillation attack" is such a loaded term that really pisses me off.

Distillation is a technical term with real meaning, and historically requires logits which Anthropic does not provide.

"Generated training data" is the correct term. It's not an "attack". And Anthropic undoubtedly also generates training data for each new generation of models, yet you never see them claim Fable is a distilled Opus.

I like this article. Rings very true.

I like Anthropic, I don't think all their talk of safety is bluff and bluster, or at least, I want to believe that the people who left OpenAI because it had lost its focus of helping humanity still want that to be their main goal. However, yes, it seems that business fears are once again causing those in charge to turn "we want to help humanity" into "we are the only ones who can help humanity, and therefore we need to be the most profitable, and the only survivors".

If you want the former ideal to survive, at Anthropic and outside of it, you need to be willing to collaborate beyond profit incentives and recouping capex. Show other labs a commitment to research and community and they will follow. Better to bring teams together rather than implicitly say you distrust them, pushing them that way instead.

"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence"

That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.

His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".

What's wrong if the roles of USA and China are reversed in technology? Why does the rest of the world care? It's not as if USA has done a great good for the world, and China has evil intentions towards the world. Infact it is the opposite in the case of AI so far.
Datapoint from the cheap end of the market: I run local models on a couple of Orange Pi SBCs and a decade-old Optiplex with no GPU. What runs usably on that class of hardware is almost entirely Chinese open weights — Qwen's MoE builds (35B total, ~3B active) are the only thing that gives me acceptable speed on CPU, with Gemma as about the only western exception. I evaluated Kimi too and ruled it out purely on size.

Whatever the strategic picture is at the top, at the bottom of the market "weights you can download and run on hardware you already own" is the whole ballgame, and right now that's mostly Alibaba's to lose.

I think releasing models for free is some 4d chess move by the chinese. Big chunk of the us stock market is fueled by ai mania, if the frontier labs turn out to be drastically less valuable than first believed, the downturm may be very bad. Think of all the big tech companies that have a ton of debt that they took to pour money into AI. It seems like a similar tactic to what Chinese car manufacturers are doing in Europe but the result may be more dramatic.
I loved this article! Regardless of how you feel about AI as an industry or tool, the economics of AI is fascinating. It's awesome to see something like this that gets into the business side a bit more.

I don't know if I agreed totally with the assessment of the risk Chinese labs pose to US labs though, in particular I think the main part I wasn't sure about was this:

> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.

How true is this? My understanding from Deepseek's original paper was that they focused heavily on optimising training and inference costs, in particular so that they can operate on cheaper (and more accessible to China) hardware.

It's possible I'm just not in the loop, but nobody seems to talk about US models innovating in this way (I'm just talking about cost-to-serve/train, not saying US AI companies don't innovate in other ways).

It seems to me at least, like there's a fair bit of evidence that AI shifting to a price based commodity market (vs a "best-model takes all" type market) would put China at a significant advantage? And even more significantly, require a pretty hefty correction of company valuations in the US?

I am more afraid of the US models to be fair. A country with no clear direction in many regards , that is threatening day in and day out the rest of the World for its own interests.
I fully agree with everything in this essay. Make distillation fair use. And let us use Mythos/Fable and Sol and successor or future models for all cybersecurity purposes.
While intelligence is said to be a replaceable commodity, oil and copper can be used in nearly the same way even if you change suppliers as long as the quality grade is matched. However, I question whether two models that produce the same benchmark answers are actually interchangeable in real world use.

Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.

Maybe not the main point of the article, but I have a doubt about the author's introduction to commoditized markets:

> - Supplier A will sell 10 units of the commodity for $20, earning $10/unit

> - Supplier B will sell 10 units of the commodity for $20, earning $5/unit

> - Supplier C will sell 5 units of the commodity for $20, earning $0/unit

> ...

> Bankruptcy risk is where fixed costs come back to the forefront: Supplier C has both fixed costs (like potentially R&D spend) and also may have taken on debt [...] It can’t price its commodity with these costs in mind — remember, the market-clearing price approximates the marginal cost of the highest-cost unit needed to satisfy demand [...]

Why can't Supplier C price their fixed costs and debt into their product? The entire reason Suppliers A and B are earning $10 and $5 per unit, and not more, is because they cannot meet demand by themselves and are therefore at the mercy of how much Supplier C is willing to charge. Couldn't Supplier C just refuse to offer 5 units of the product at a price that would bankrupt them?

Sincerely, an interested observer of business/economics.

I disagree with the first half quite a bit, COGS ultimately depends on the use case. If someone just wants something that a smaller model can do, running a local model on phone is going to have a negligible cost close to running any other piece of software. The alternatives to running a model also determines COGS, and even Jensen Huang has distinguished between the job and the work for the job that AI is capable of doing. Smaller models are always going to win in efficiency too.

I also heavily disagree with this no-marginal cost in software distribution view whenever I see it, bit rot is real, and someone is paying a marginal cost whenever they do an update. You have to re-distribute with changes whenever anything changes. These costs are just hidden because things are ad-supported or bundled in some way. These costs are also kept low because of standards and open source, but could become high anytime. Additional licensing also has costs.

That said, I couldn't agree more with the last paragraph, charging a high price for models would be better than denying access for any model that wants to stay relevant.

The article makes a great point that the token industry is going to be commoditized as time goes on.

Following this argument the key for each player will be the underlying cost structure and serving capacity to offset the upfront R&D cost.

The cost infrastructure will be driven by access to cheap electricity and cheap chips. The capacity will be driven primarily by depth of pockets now to buy all available supply in chips/mem/data center building capacity. While China is certainly in the lead on cheap energy, I am wondering if they can/want to beat the > 1tn USD being spent on data centers right now. Following the example in the article:

If company C from China sells 10 units for 20 USD produced for 10 USD they pocket 100 USD.

If company A from America can sell 100 units for 20 USD produced for 15 units, they pocket 500 USD or 5/6th of the market's profits.

> By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage.

I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].

If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.

But that's a big if we just don't know for sure.

1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...

I'm worried that any ban on Chinese AI models might be an excuse to get mass surveillance.
No one should be afraid of anything. Fear is a terrible advisor. Keep your eyes open, try to read the context as careful as you can and adapt as best as you can. Don’t spent too much time trying to be an oracle, never works out…
You couldn't be more wrong. What's being sold are chunks of time with access to specialized hardware resources. Through which model or with which device is irrelevant; the one winning, and continuing to win for some time, is Nvidia, and that company is American. As long as they have the H100, B200, or B300, China won't be able to compete with the American strategy, no matter how many new models they release, because these types of cards require incredibly powerful hardware to run.
"Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute."

It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.

The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.