It’s great to buy dollars for a penny, but the guy selling em is going to want to charge a dollar eventually…
Do you feel there is enough visibility and stability around the "Prompt -> API token usage" connection to make a reliable estimate as to what using the API may end up costing?
Personally, it feels like paying for Netflix based on "data usage" without having anyway for me to know ahead of time how much data any given episode or movie will end up using, because Netflix is constantly changing the quality/compression/etc on the fly.
I agree that ex ante it’s tough, and they could benefit from some mode of estimation.
Perhaps we can give tasks sizes, like T shirts? Or a group of claudes can spend the first 1M tokens assigning point values to the prospective tasks?
Take the response on another post about Claude Code.
https://news.ycombinator.com/item?id=47664442
This reads like even if you had a rough idea today about what usage might look like, a change deployed tomorrow could have a major impact on usage. And you wouldn't know it until after you were already using it.
Of course, I have no idea how MS is justifying the Copilot pricing. I can't imagine any world in which it is sustainable, so I'm trying to get as much as I can out of it now before they jack up prices.
Now we’re going to find out what these tools are really worth.
So I noticed the model is purposefully coming with dumb ideas or running around in circles and only when you tell it that they are trying to defraud you, they suddenly come back with a right solution.
It works out even if some customers are able to eat a lot, because people on average have a certain limit. The limits of computers are much higher.
Pay-per-token is really the only way it can work. If some kind of fixed monthly price is desirable, then there should be a quota the user can assign, and then the agent could e.g. slow itself down by 50% when 50% of the quota is spent, another 50% at 75%, etc, to make it last longer..
As a side thought, I wonder how it could affect an agent's behavior if the information of this token usage/limit was brought to it..
I wonder if there could be something like that, maybe even a progressive rate limiting, where after a certain number of tokens or another metric of use, then the speed slows down a LOT.
Not saying that I would love that as a consumer, as I'd prefer this all-you-can eat, unlimited data plan, but I wonder if that would be a compromise that could work, as it seems to have worked OK with the telecom space.
edit: the nerd in me loves the irony of me making the above comment and then later seeing your username as flux :-)
If an hour of an excellent developer's time is worth $X, isn't that the upper bound of what the AI companies can charge? If hiring a person is better value than paying for an AI, then do that.
They can charge whatever they want, I think many people like to make business decisions based on relative predictability or at least be more aware that there's a risk. If they want it to be "some weeks you have lots of usage, some weeks less, and it depends on X factors, or even random factors" then people could make a more informed choice. I think now it's basically incredibly vague and that works while it's relatively predictable, and starts to fail when it's not, for those that wanted the implied predictability.
I'm not sure how businesses budget for llm APIs, as they seem wildly unpredictable to me and super expensive, but maybe I'm missing something about it.