back
257 comments
> One user, who asked not to be identified, said it has been impossible to advance his project since the usage limits came into effect.

Vibe limit reached. Gotta start doing some thinking.

Who would have though including a hard depedency on third part service with unclear long term availability would be a problem!

Paid compilers and remotely acessible mainframes all over again - people apparently never learn.

Right, but these companies are selling their products on the basis that you can offload a good amount of the thinking. And it seems a good deal of investment in AI is also based on this premise. I don't disagree with you, but it's sorta fucked that so much money has been pumped into this and that markets seem to still be okay with it all.
> Gotta start doing some thinking.

The fact they declared their own project as “impossible to advance” given the situation reveals they are unwilling to go the thinking route right now.

Hmmm, I am 99% sure the users are not vibe coders who can't code, those are on tools like lovable, not messing with terminal tools.
More than some thinking. They’ll probably need to think hardest or even ultrathink to keep the project moving forward.
Came to comment on the same quote.

I'm surprised, but know I shouldn't be, that we're at this point already.

I honestly feel sorry for these vibe coders. I'm loving AI in a similar way that I loved google or IDE magic. This seems like a far worst version of those developers that tried to build an entire app with Eclipse or Visual Studio GUI drag and drop from the late 90s
He did not pass the vibe check.
The funny thing is Claude 4.0 isn't even that 'smart' from a raw intelligence perspective compared to the other flagship models.

They've just done the work to tailor it specifically for proper tool using during coding. Once other models catch up, they will not be able to be so stingy on limits.

Google has the advantage here given they're running on their own silicon; can optimize for it; and have nearly unlimited cashflows they can burn.

I find it amusing nobody here in the comments can understand the scaling laws of compute. It seems like people have a mental model of Uber burned into their head thinking that at some point the price has to go up. AI is not human labor.

Over time the price of compute will fall, not rise. Losing money in the short term betting this will happen is not a dumb strategy given it's the most likely scenario.

I know everybody really wants this bubble to pop so they can make themselves feel smart for "calling it" (and feel less jealous of the people who got in early) and I'm sure there will be a pop, but in the long term this is all correct.

I played with Claude Code using the basic $20/month plan for a toy side project.

I couldn't believe how many requests I could get in. I wasn't using this full-time for an entire workweek, but I thought for sure I'd be running into the $20/month limits quickly. Yet I never did.

To be fair, I spent a lot of time cleaning up after the AI and manually coding things it couldn't figure out. It still seemed like an incredible number of tokens were being processed. I don't have concrete numbers, but it felt like I was easily getting $10-20 worth of tokens (compared to raw API prices) out of it every single day.

My guess is that they left the limits extremely generous for a while to promote adoption, and now they're tightening them up because it’s starting to overwhelm their capacity.

I can't imagine how much vibe coding you'd have to be doing to hit the limits on the $200/month plan like this article, though.

They're likely burning money so I can't be pissed off yet, but we see the same Cursor as well; the pricing is not transparent.

I'm paying for Max, and when I use the tooling to calculate the spend returned by the API, I can see it's almost $1k! I have no idea how much quota I have left until the next block. The pricing returned by the API doesn't make any sense.

That’s funny I literally started the $200/month plan this week because I routinely spend $300+/month on API tokens.

And I was thinking to myself, “How does this make any sense financially for Anthropic to let me have all of this for $200/month?”

And then I kept getting hit with those overloaded api errors so I canceled my plan and went back to API tokens.

I still have no idea what they’re doing over there but I’ll happily pay for access. Just stop dangling that damn $200/month in my face if you’re not going to honor it with reasonable access.

I need to see a video of what people are doing to hit the max limits regularly.

I find sonnet really useful for coding but I never even hit basic limits. at $20/mo. Writing specs, coming up with documentation, doing wrote tasks for which many examples exist in the database. Iterate on particular services etc.

Are these max users having it write the whole codebase w/ rewrites? Isn't it often just faster to fix small things I find incorrect than type up why I think it's wrong in English and have it do a whole big round trip?

I'm not sure this is "intentional" per se or just massively overloaded servers because of unexpected demand growth and they are cutting rate limits until they can scale up more. This may become permanent/worse if the demand keeps outstripping their ability to scale.

I'd be extremely surprised if Anthropic picked now of all times to decide on COGS optimisation. They potentially can take a significant slice of the entire DevTools market with the growth they are seeing, seems short sighted to me to nerf that when they have oodles of cash in bank and no doubt people hammering at their door to throw more cash at them.

The other day I was doing major refactorings on two projects simultaneously while doing design work for two other projects. It occurred to me to check my API usage for Gemini and I had spent $200 that day already.

Users are no doubt working these things even harder than I am. There's no way they can be profitable at $200 a month with unlimited usage.

I think we're going to evolve into a system that intelligently allocates tasks based on cost. I think that's part of what openrouter is trying to do, but it's going to require a lot of context information to do the routing correctly.

> One user, who asked not to be identified, said it has been impossible to advance his project since the usage limits came into effect. “It just stopped the ability to make progress,” the user told TechCrunch. “I tried Gemini and Kimi, but there’s really nothing else that’s competitive with the capability set of Claude Code right now.”

PMF.

This is what really makes me sceptical of these tools. I've tried Claude Code and it does save some time even if I find the process boring and unappealing. But as much as I hate typing, my keyboard is mine and isn't just going to disappear one day, have its price hiked or refuse to work after 1000 lines. I would hate to get used to these tools then find I don't have them any more. I'm all for cutting down on typing but I'll wait until I can run things entirely locally.
I made a quick site so you can see what tools are using the most context and help control it, totally free and in your browser.

https://claude-code-analysis.pages.dev/

“It just stopped the ability to make progress,” the user told TechCrunch. “I tried Gemini and Kimi, but there’s really nothing else that’s competitive with the capability set of Claude Code right now.”

This is probably another marketing stunt. Turn off the flow of cocaine and have users find out how addicted they are. And they'll pay for the purest cocaine, not for second grade.

I wish models which we can self-host at home would start catching up. Relying on hosted providers like this is a huge risk, as can be seen in this case.

I just worry that there’s little incentive for bit corporations to research optimising the “running queries for a single user in a consumer GPU” use case. I wonder if getting funding for such research is even viable at all.

It’s like you buy an M4 MacBook and silently Apple throttles it and makes it an M1. If that happened every tech magazine would write about it, consumer advocates would be in rage.

How’s it possible that AI companies sell you a product for 100 USD a month and silently degrade it?

I have the $100 plan and now quickly get downgraded to Sonnet. But so far have not hit any other limits. I use it more on the weekends over several hours, so lets see what this weekend has in store.

I suspected that something like this might happen, where the demand will outstrip the supply and squeeze small players out. I still think demand is in its infancy and that many of us will be forced to pay a lot more. Unless of course there are breakthroughs. At work I recently switched to non-reasoning models because I find I get more work done and the quality is good enough. The queue to use Sonnet 3.7 and 4.0 is too long. Maybe the tools will improve reduce token count, e.g. a token reducing step (and maybe this already exists).

I was wondering when we'd get to the end of the reduced fare portion of the bubble. With the layoffs, increase in prices, we're going to see a lot more vibe coding to justify the price of ai.

With the talented employees laid off, I predict there will be some VERY LARGE code mistake pushed, either banking or travel, and we'll either see the burst, or they'll move on to strictly enterprise/gov opportunities. (see: grok for gov)

>Super Free [Introduction, lots of free options] >Reduced Fare [higher prices, smaller pool of free options](We are here) >Premium Only [only paid options, mostly whales]

Does anyone not realize they are just using the typical drug dealer type business model? I used to do cocaine and it was a similar vibe.

They will turn you into an AI junkie who no longer has motivation to do anything difficult on your own (despite having the skills and knowing how), and then, they will dramatically cut your usage limit and say you’ll need to pay more to use their AI.

And you will gladly pay more, because hey you are getting paid a lot and it’s only a few hundred extra. And look at all the time you save!

Soon you’re paying $2k a month on AI.

Is it really worth it to use opus vs. sonnet? sonnet is pretty good on its own.
Id like to hear about the tools and use cases that lead people to hit these limits. How many sub-agents are they spawning? How are they monitoring them?
oh yea looks like everyone and their grandma is hitting claude code

https://github.com/anthropics/claude-code/issues/3572

Inside info is they are using their servers to prioritize training for sonnet 4.5 to launch at the same time as xAI dedicated coding model. xAI coding logic is very close to sonnet 4 and has anthropic scrambling. xAI sucks at making designs but codes really well.

the day of COGS reckoning for the "AI" industry is approaching fast
All you people who were happy to pay $100 and $200 a month have ruined it for the rest of us!!
Yeah, i noticed it falls back to sonnet much quicker too. Most days within the first few minutes.

And the 529.. it's borderline unusable at times.

But the worst part: While claude code, the tools and cli etc, has become much better over the last weeks; it seems the models or prompts have gotten worse. It will do things like add a test, see that it fails, claim it was already broken and out of scope. Or maybe i ask it to implement this, DO NOT use Y. Use Y anyway. Sometimes i ask it to update a test & it decides to revert all changes in the application code.

I am very close to cancelling my plan again. Maybe it is a great deal when comparing too api pricing, but that is not a fair comparison; prepaid/pay-as-you-go is always way more expensive. And if they priced it wrong, change the pricing? Degrading a new service you are developing to save cost is even more stupid then the burn money until ??? profit stats imo hehe

Hardly surprising.

AWS Bedrock which seems to be a popular way to get access to Claude etc. while not having to go through another "cloud security audit", will easily run up ~20-30$ bills in half-hour with something like Cline.

Anthropic likely is making bank with this and can afford to lose the less-profitable (or even loss-making) business of lone-man developers.

So far I’ve had 3-4 Claude code instances constantly working 8-12 hours a day every day. I use it like a stick shift though. When I need a big plan doc, switch to recommended model between opus and sonnet. And for coding, use sonnet. Sometimes I hit the opus limit but I simply switch to sonnet for the day and watch it more closely.
I guess flat fee AI subscriptions are not a thing that is going to work out.

Probably better to stay on usage based pricing, and just accept that every API call will be charged to your account.

I don't think CLI/terminal-based approaches are going to win out in the long run compared to visual IDEs like Cursor but I think Anthropic has something good with Claude Code and I've been loving it lately (after using only Cursor for a while.) Wouldn't be surprised if they end up purchasing Cursor after squeezing them out via pricing and then merging Cursor + Claude Code so you have the best of both worlds under one name.
I think that Cursor is doing the same. A couple of weeks ago they removed the 500 prime model requests limit per month in the $20 plan, it seemed like this was going to be good for users, in fact it's worse, my impression is that now the limit is effectively much lower, and you can't check anymore in your account's dashboard how many of these requests you've made over the last month.
That must somehow be illegal, at least in the consumer space. I have noticed quicker degrade to sonnet, but don't often hit limits ($200 plan). Seems none of these guys can afford loyalty, so I will be skipping between 'tool of the month' instead of sticking with one. New companies with new VC money are good for a few months and then degrade, so it's not hard to do.
Customers are always chasing the next big thing in this space. As a programmer who’s worked on mobile UIs and backends using CC, I can say the appeal isn’t really the model itself—it’s the form factor. It’s a decent product, but hardly groundbreaking or essential.
There’s been a ton of ‘service overloaded’ errors this week so it makes sense that they had to adjust it.

Personally I’ve never hit a usage limit on the $100 plan even when running several Claude tabs at once. I can’t imagine how people can max out the $200 plan.

I went from pro to max because I hve been hitting limits, I could tell they were reducing it because I used to go multiple hours on pro but now its like 3. Congrats Anthropic you got $100 more out of me, at the cost of irrecoverable goodwill
Ha was just talking about this coming down the pipeline with folks days ago (in so many words) https://news.ycombinator.com/context?id=44565481
As a independent dev, I haven't had the need to pay for any AI yet. When I run into my limit at one company, I switch to the next one. Not always the same experience, but the next day I can start fresh again.
They didn’t reduce it they actually increased it. I was using it for 14 hours straight without any issues. I think they did that to stay competitive, but now it seems like it's back to normal.
What are these people doing to hit their limit this fast?

I put $30 on it, use it daily at work, and I still have half of it left. Are Zed agents just that optimized? I doubt it

opencode with kimi-k2 is my backup just in case claude is down or I hit the limits on the max 20x plan.
I think it was just an outage that unfortunately returned 429 errors instead of something else.