back

by kristjansson·1y ago·view on hn ↗
This is ridiculous. Since the actual training code for DeepSeek is _not_ public, this is a based only on the technical report, which mentions PTX one (1) time in §3.2.2 Efficient Implementation of Cross-Node All-to-All Communication:

> Specifically, we employ customized PTX (Parallel Thread Execution) instructions and auto-tune the communication chunk size, which significantly reduces the use of the L2 cache and the interference to other SMs.

So they have some intrinsic in some part of their training framework. That's it.

1 comments
It all feels like an attempt to prevent replication to further tank the market. Not necessarily the technical details, but the reporting thereo, which was spammed all over Reddit and HN
This whole episode is weird. I can’t tell how much of the popular reporting is misinformed and how much has been disinformed. R1 (and sorta V3) are clearly progress, but are definitely not step-function improvements to prior SOTAs.
They are definitely step function improvements in regard to inferrance cost of the same level of reasoning intelligence.
Only because OpenAI overpriced o1, like they did with GPT-4.