back

by 7777777phil·5mo ago·view on hn ↗
$730B pre-money for a company where each model is roughly 2x profitable on its own, but each next model costs 10x the last. The whole thing only works if scaling keeps delivering. Research (Sara Hooker et al.) is not encouraging on that front, compact models already outperform massive predecessors on downstream tasks while scaling laws only predict pre-training loss reliably.

Wrote about both the per-model math and the scaling question:

(1) https://philippdubach.com/posts/ai-models-as-standalone-pls/

(2) https://philippdubach.com/posts/the-most-expensive-assumptio...

EDIT: Removed the dot after et; bc apparently it's an entire word (the more you know..)

6 comments
> each model is roughly 2x profitable on its own, but each next model costs 10x the last. The whole thing only works if scaling keeps delivering.

This is a decent argument, but it's not the death knell you think.

Models are getting 99% more efficient every 3 years - to get the same amount of output, combined with hardware and (mostly) software upgrades - you can use 99% less power.

The number of applications where AI is already "good enough" keeps growing every day. If the cost goes down 99% every three years, it doesn't take long until you can make a ton of money on those applications.

If AI stopped progressing today, it would take probably a decade or longer for us to take full advantage of it. So there is tons of forward looking revenue that isn't counted yet.

For the foreseeable future, there are MANY MANY uses of models where a company would not want to host its own models and would be GLAD to pay an 4-5x cost for someone else to host the model and hardware for them.

I'm as bullish on OpenAI being "worth" $730B as I was on Snap being worth what it IPO'd for - which it's still down about 80% (AFTER inflation, or about ~95% adjusting for gold inflation).

But guess what - these are MINIMUM valuations based on 50-80% margins - i.e. they're really getting about ~$30B - the rest is market value of hardware and hosting. OpenAI could be worth 80% less, and they could still make a metric fuck-ton of money selling at IPO with a $1T+ market cap to speculative morons easily...

Realistically, very rich people with high risk tolerance are saying that they think OpenAI has a MINIMUM value of ~$100B. That seems very reasonable given the risk tolerance and wealth.

When models get cheaper to run for OpenAI, they also get cheaper for everyone else. It gets commoditized. AI might be able to do more, but most people aren’t going to pay for a thing they could get for free. See the many models on Huggingface as examples of that.

And as the number of things AI is “good enough” at increases, the list of things on the frontier that people will want to pay OpenAI for shrinks. Even if OpenAI can consistently churn out PhD level math, most companies don’t care about that.

So a necessary (but not sufficient) condition for the math to work out is that frontier tasks still exist and are profitable. This is why CEOs keep hyping up AGI. But what they really want is for developers to keep paying to get AI to center a div.

> get cheaper to run

Irrelevant. The model is the moat

> most companies don’t care about that.

Wrong. They will use the model that gives them an edge. If they are using a PhD but their competitors are using Einstein, they will lose.

> center a div

For sure a common use case, but is bot what the CEO is concerned about with AI.

> Wrong. They will use the model that gives them an edge. If they are using a PhD but their competitors are using Einstein, they will lose.

For some tasks that matters. But for a lot of tasks, "good enough but cheaper" will win out.

I'm sure there will be a market for whichever company has the best model, but just like most companies don't hire many PhD's, most companies won't feel a need for the highest end models either, above a certain level.

E.g. with the release of Sonnet 4.6, I switched a lot of my processes from Opus to Sonnet, because Sonnet 4.6 is good enough, and it means I can do more for less.

But I'm also experimenting with Kimi, Qwen, Deepseek, and others for a number of tasks, including fine-grained switching and interleaving. E.g. have a cheap but dumb model filter data or take over when a sub-task is simple enough, in order to have the smart model do less, for example.

Models will get smarter and cheaper. For those that are burned directly into silicon, there will be a market for old models - as the alternative is to dump that silicon in a landfill.

For models that run on general-purpose AI hardware, I don't know why the vendors would waste that resource on old models.

Who says anything about old models? What we're seeing is that as the frontier models get better, we get cheaper, better small models that leverage the advanced but cost a fraction. At the same time, hardware provides morez cheaper options. Sometimes far faster options too (e.g. Cerebras).

In terms of price, I can get 1m output tokens from Deepseek for 40 cents vs. 25 dollars for Opus, and a number of models near the 1-2 dollar mark that are increasingly viable for a larger set of applications.

Providers will keep running those cheaper models as long as there's demand.

Larger models need more hardware resources to run

And, depending on effort settings, they do more 'thinking', i.e., use more rounds of inference to generate longer internal chains of thought

Both very good reasons to prefer a smaller model, if the small model is good enough for the task

> The model is the moat

What model? GPT4o certainly isn’t a moat for open ai. They need to keep training better and better models because qwen3, kimi k2.5 etc constantly nipping at their heels.

> Wrong. They will use the model that gives them an edge. If they are using a PhD but their competitors are using Einstein, they will lose.

It depends on the business. As much as I’d love to engage a PhD or an Einstein in my Verizon customer support call, it isn’t going to net the call center any value to pay for that extra compute.

It's a moat. Yes, they must keep refilling it, but it's all they have.

My PhD vs Einstein analogy was bad. What I mean is stupid vs smart. Nobody is going to pay for a stupid model when they can pay a bit more for smart.

But what if all models that are smart enough for the task? Then its about price no?
> If they are using a PhD but their competitors are using

god what are these assumptions

More analogy than assumption. And admittedly a poor analogy.
> Models are getting 99% more efficient every 3 years - to get the same amount of output, combined with hardware and (mostly) software upgrades - you can use 99% less power.

Even if true, this still doesn't bend the curve when paying for the next model.

> If AI stopped progressing today, it would take probably a decade or longer for us to take full advantage of it. So there is tons of forward looking revenue that isn't counted yet.

If this is true, it's true for the technology overall, and not necessarily OpenAI since inference would get commoditized quickly at that point. OpenAI could continue to have a capital advantage as a public stock, but I don't think it would if the music stopped.

I would actually like to see the real math currently.

The market adoption has increased a lot. The cost to serve has come down a lot per token.

Model sizes have not increased exponentially recently (The high point being the aborted GPT-4.5), most refinement recently seems to be extending training on relatively smaller models.

When you take this into account together, the relative training to inference income/cost ratio likely has actually changed dramatically.

I love that you are already confident fitting a curve. I want some of that swagger in my life.
I was thinking the same thing.
"If AI stopped progressing today, it would take probably a decade or longer for us to take full advantage of it."

AI stopped progressing, or LLMs? I really dislike people throwing the term AI around.

For the purposes of their argument, I don’t think the distinction matters.
> Models are getting 99% more efficient every 3 years

The LLM industry has only be around for like 4 years. Extrapolating trends from that is pretty naive.

> 99% more efficient every 3 years

It's 2x efficiency. Then I'd take 50% less power instead of ridiculous 99% less power.

GPT-4 came out 3 years ago and you can run comparable models for 1% of the cost nowadays. That is not 2x efficiency. That's two orders of magnitude in end-to-end compute efficiency.
you're looking at nearly the entire curve of the tech's development. that's like saying lightbulbs became 99% more energy efficient and therefore will become another 99% more energy efficient. but most techs follow an S curve.
But S curves are boring and dont moon
>you're looking at nearly the entire curve of the tech's development

That's a pretty strong statement that would need some data or at least a mathematical argument to back it up. Otherwise it's like saying in the 1980s that PCs with 640kB RAM have reached their pinnacle in terms of what users can expect in real life benefits and there's no reason to keep pushing the tech.

*entire curve to-date (I should have clarified). Yes it will get better for a long time, but where we are on the curve is harder to say. Lots of metrics to choose from, like "well it's incorrect 90% less often than a year ago, so that's a 10x improvement!". But the real metric that matters is how useful it is to people, and based on user data it looks like the only area it's getting exponentially more useful YoY is for programming. Lot of coders using it 10x more than before to code 10x faster. Not sure any other profession uses it for more than a juiced-up search engine / proofreader.
Tbf that sounds like a strong bias from someone who works exclusively in software development and simply hasn't found other uses. But I have worked with integrating LLMs across quite a few applications and departments by now and I can comfortably say that programming is not the only thing where we see extreme benefits. I wouldn't even say it's the area that has seen the most benefit so far. There used to be a lot of mundane work outside of software development that was easy prey even for early models. And with the current cutting edge models I'm pretty sure that you could replace >75% white collar jobs if you just get the context engineering right. That's the hard part right now, not the raw intelligence necessary for arbitrary data processing. But frameworks are getting there fast.
> most techs follow an S curve.

All techs, eventually.

How do we know how much it costs? Or is this just based off the token pricing?
That's the bingo of the question... The entire argument is token pricing, which can be subsidized.
We said all the same shit about VR, dude. Even had a global pandemic show up to boost everyone's interest in the key market of telepresence. Turns out the merry go round can stop abruptly.
No. Like many of us, I never saw much value in VR. LLMs have undeniable value that is general and broad. Now, does that mean OpenAI has a moat? No, it does not.
We also said that about VR.
Did we?! You and Mark Zuckerberg maybe.
"Am I nothing to you?" --Tim Cook
ok, but everything depends on your numbers being correct. 99% improved efficiency seems kind of a way too optimistic prediction.
> Models are getting 99% more efficient every 3 years - to get the same amount of output, combined with hardware and (mostly) software upgrades - you can use 99% less power.

This is such a poor argument for a number of reasons.

1. Three years ago is basically when the "AI race" really kicked off amongst the frontier companies. You're effectively comparing a car from the 1920/30's to a modern car.

2. Past performance is not an indicator of future performance. You can't just say that LLM's will grow and improve at a fixed rate for all time, that isn't how they or anything else works in the real world.

3. Since it's an open secret that companies like Anthropic and OpenAI are running their models at a loss, a static 99% cheaper every three years arc still puts these companies at a net negative position unless compute, energy and water all somehow start getting 99% cheaper every three years.

> Models are getting 99% more efficient every 3 years

How many years total are you basing this on?

Obligatory XKCD: https://xkcd.com/605/
> Models are getting 99% more efficient every 3 years

Ugh. Someone has to do this: https://xkcd.com/605/

> each next model costs 10x the last

Yes, but there's a chance that actually training is done more or less for free by companies like OpenAI. The reason being that they do a gigantic amount of inference for end users (for which they get paid), but their servers can't be constantly utilized at 100% by inference. So, if they know how to schedule things correctly (and they probably do), they can do the training of their new model on the unutilized compute capacity. If you or I were to pay for that training, it would be billions of dollars, but for them it is just using compute that otherwise would be idle.

What makes you think this trend will continue? In a situation with finite resources (eg the number of parameters), the default is to assume things will plateau.
I was reading a paper on dark silicon and how it broke the beautiful scaling laws of the past (Moore's law/Dennard Scaling). We hit a wall, innovated and at the moment, the hardware industry is thriving. To me, that means scaling the industry and riding that momentum wasn't wrong. In fact, it allowed us to be where we are today.

Why are we so against, in principle, to the current pre-training scaling laws? Perhaps, we'll require new innovations at some point, but the momentum allows us to reach to newer heights that we've never climbed before.

> EDIT: Removed the dot after et; bc apparently it's an entire word (the more you know..)

From latin "et alia", abbreviated as "et al." - it's not a single word but an expression.

Et is an entire word and doesn’t need a period at the end.