tokens are going to be as cheap as electricity within the decade

353 points•teoruiz•8 days ago•227 comments•

227 comments

jetrink8 days ago
> Tokens become cheaper than tool calls

The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep. I think this is a good time to invoke Stein's Law: "If something cannot go on forever, it will stop." These efficiency improvements won't continue forever. It's more likely that the per-call cost of high-quality, compiled software like grep will be a lower-bound that LLMs asymptotically approach, rather than a line that they blow past with perpetual exponential progress. (Barring a true breakthrough in something like quantum computing or room-temperature superconductors.)

sanderjd8 days ago
Yeah I bumped on that too. If it's possible to make llms cheaper than current grep, then it is also almost certainly possible to make grep cheaper.
serbuvlad8 days ago
You can burn anything* into an ASIC to make it cheaper per-call.

non-backreferencing grep is not very difficult to implement in an ASIC either. But it's probably not worth it because of how relatively rarely you use it and of the data transfer costs.

LLMs are great candidates for ASIC-burning because they're slow compared even to network speeds and run all the time. The issue is that you don't want to burn a specific model or architecture that then becomes obsolete.

So you've got two possible futures, and both guarantee large price drops: (a) LLMs keep getting better and better and better, so ability/$ keeps rising; or (b) LLMs plateau in ability, in which they will start getting ASIC'd.

scotty798 days ago
depends on what you are grepping ... greapping a large file might be more expensive one day than generating n-th token with LLM that works fully in hardware

you could make hardware implementation of grep and store the file itself next to it in some ROM but that's not a very useful grep ... while hardware LLM is exactly as useful as software LLM only orders of magnitude faster

rdsubhas7 days ago
grep is deterministic. Llm is probabilistic. Llm can be transferred to a tiny quantum cpu or a lower precision float.

When the author wrote Llm can be as cheap as a tool, I read it as not equivalent. They even said the Llm can be embedded into a tool.

Their point was, the higher level use case — like classification — could become as cheap as grep. Which is quite well possible.

BearOso7 days ago
4-5 orders of magnitude is huge. Assuming an order of base 10, it's 10000x-100000x. So a call to grep may return in 1s on a typical PC. That means a GPT call takes equivalent energy of 10000-100000 PCs to do the same in 1s. That's a difference that can't be equalized with scaling. It would require a revolutionary breakthrough.

I also don't understand where the idea that frontier models are getting better efficiency comes from. The results are certainly improving, but that comes from feedback and multiplexing requests, which cost more.

chewbacha7 days ago
Recall the articles claim that efficiency is gaining 2.5 orders of magnitude per year.

Something seems off about this.

FranOntanaya8 days ago
LLM is spicy memoizing, so it can potentially be faster than a tool call. But people will spend a month tweaking and testing to ensure they have the level of determinism they need, which means it's more expensive, and that they should have used actual memoization in the first place.
lelanthran8 days ago
> The author observes that a call to GPT-5.6 Luna is only 4-5 orders of magnitude more expensive than grep, and then predicts that at current rates of progress, calling an LLM will soon be cheaper than a grep.

At some future point where LLM hardware is cheaper than simply running grep, then grep equivalent would benefit from those selfsame hardware improvements and be cheaper to run as well, probably still by the same ratio.

jandrese7 days ago
A classic case of someone projecting out to infinity from just after the first bend of the S curve.
cs7028 days ago
I found the OP insightful and worth a read. Thank you for sharing it on HN.

The only aspect that is poorly analyzed by the OP is business model viability. All players are investing insane amounts of money in infrastructure with the expectation that their future profits will justify all that investment. The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."

The OP glosses over questions of business model viability with a brief qualitative discussion and very little hard data. For example, to earn an annual return > 10% on every trillion dollars of capital sunk into infrastructure, the owners of that infrastructure must earn free cash flow (operating profit less investment) in excess of $100 billion per year in perpetuity. Is that feasible? Why? How?

The OP does not really consider such questions.

sanderjd8 days ago
I think the article's analysis is basically right in a vacuum. That is, I think it's clear that inference is a viable business model. But what isn't clear is whether it will be such a profitable business model for any given company that it will justify the investment that company has taken. I kind of think the winners might be a follow-on generation of companies that focus on this commodity inference business model instead of the invent-machine-god-first "business model" and thus are wiser about their level of investment and capital costs.
cs7028 days ago
> I think it's clear that inference is a viable business model.

You may be right. I'm not so sure. Inference looks like a viable business model for those operators that have SOTA infrastructure in place, but the investment required to have it is enormous, and appears to be never-ending, because if an operator stops investing aggressively, its infrastructure quickly becomes non-competitive, and customers will quickly leave for alternatives. SOTA infrastructure is a moving target.

KaiserPro7 days ago
> I think it's clear that inference is a viable business model.

Only if you also have the model thats better than anyone else's.

As soon as models are free, or there are no newer models (assuming thats going to happen, and thats not a given) then the only thing you can compete on is price.

This means that the only thing you have to differentiate is either price, speed or ease of use. (or regulatory capture...)

We are at pets.com level of spend currently. Unless model development becomes cheaper, then we are going to run out of novel debt but not really debt mechanisms.

twoodfin8 days ago
A phrase comes to mind: "Your margin is my opportunity."
otherme1238 days ago
The problem source is that the "cost" of tokens are taken at face value from business that are losing money at record speeds. E.g. https://artificialanalysis.ai says "doing task A costed us $10 using OpenAI", and that is the "cost" the OP used as basis for "tokens are cheap". Meanwhile OpenAI is losing $19 for each $1 in revenue... So right now OpenAI should be charging around $200 to do task A just to break even, but that would mean their use base would collapse.
pixl977 days ago
There are two different things occurring here.

One is how much does it cost OpenAI to train the model.

The other is, if I stole OpenAI's model how much would it cost for me to run it?

R&D costs versus operational costs. Operational costs are very likely profitable. R&D is catastrophically expensive currently.

asdff7 days ago
>The winner or winners in the AGI race, they believe, will find the proverbial "pot of gold at the end of the rainbow."

The hubris of this is really astounding too. There is no technology out there that some company develops and has not been reverse engineered and copied and manufactured at scale by competitors before long. You can't stop this from happening. People will leave the company or be poached and proliferate what they have built in the past. Every country that wanted a nuke has a nuke, after all.

The only way to keep the secret fire from leaking out would be to have AGI's first move be to lock the doors and prevent anyone from ever leaving company property again.

hattmall7 days ago
Google's search has had a long run, and seems much less complex than LLMs.
tim3337 days ago
There's some alternative analysis here that the current economics aren't too bad https://x.com/ramez/status/2102506182279573962

Also another take on AI costs falling https://x.com/EpochAIResearch/status/2102510281176023529 Methodology is how much it costs to do the same task a while later. Gives fallen ~47% a quarter.

alpineidyll38 days ago
Exactly. "Cost-to-distill" is a critical parameter. Right now usage of frontier models for all tasks is both subsidized and irrationally popular even at the subsidized price. Deepseek would solve most tasks faster and 10x cheaper. I agree with the author that just as Deloitte exists, frontier labs will exist. But not because their products are proprietary technical marvels or gods, but rather because of branding.
ido8 days ago
DS wouldn't be 10x cheaper than the subsidized subscription plans from openai/anthropic. Although it is of course much cheaper than the enterprier/API pricing- I think if you're on the subscription plans, you can't beat that on performance per price.
abirch8 days ago
Too Cheap to Meter reminds me of the promise of Nuclear Power in 1954

"It is not too much to expect that our children will enjoy in their homes electrical energy too cheap to meter,..." Lewis Strauss

https://en.wikipedia.org/wiki/Too_cheap_to_meter#Origins

Oddly enough my power bill was metered and big.

qlte8 days ago
It's a deliberate reference/meme that is basically used to acknowledge the precedent of overly exuberant predictions of cost in an emerging technology but argue "however, this time it's true".

Of course, perilous territory for future irony depending on how your prediction plays out.

m4637 days ago
I remember up to around the time the PC came out, you were billed for computer time.

Later, you were billed for time connected to the "internet" (compuserve or aol or whatever)

Around when the iphone came out, software went from tens or hundreds of dollars to pennies, then free.

On the other hand legal advice has always been expensive, because a good answer is worth it.

Medical advice is worth it. Investing advice is worth it.

(That said, I wonder if with home solar and batteries if electricity will ever "generally" go down in price to normal people)

etatester8 days ago
It's really quite unfortunate that the promise was not delivered, mostly for political reasons. I hope that a new wave of reactors and the dire need for clean energy restarts the nuclear race.
0cf8612b2e1e8 days ago
The economics are not there. Solar + batteries are much cheaper per watt today and still improving. Even better, the solar can come online instantly and expand while US nuclear takes twenty years to start generating any energy.
FearNotDaniel7 days ago
One excellent nerd-sniping side-effect of those political reasons is that the good folks of Austria built an entire nuclear power station and then never put it into service [0]. Open for guided tours on Fridays [1].

[0] https://en.wikipedia.org/wiki/Zwentendorf_Nuclear_Power_Plan...

[1] https://windows2008.zwentendorf.com/en/location-npp/

wat100008 days ago
Nuclear was never going to give us too cheap to meter regardless of politics. Uranium just isn’t that cheap. Fuel costs are lower then coal or gas, but not so low that operators would just not bother to charge for it
wat100008 days ago
On the other hand, this did work out in other areas. I pay a flat monthly rate for all-I-care-to-eat internet access, for example. My email provider has limits on storage space but I don’t get charged per email sent or received. It’s not a crazy concept on its face, nuclear power just didn’t work out as well as was hoped.
bergie7 days ago
Having your own solar can be "too cheap to meter", though it is definitely a feast and a famine kind of a thing.

We had batteries full and were spending electricity on all kinds of luxury things like whole-day internet and desalination for several weeks, and now that it has rained for over a week we're starting to turn non-critical systems off to keep the lights on.

initramfs8 days ago
https://www.wired.com/2008/02/ff-free/

"Between digital economics and the wholesale embrace of King's Gillette's experiment in price shifting, we are entering an era when free will be seen as the norm, not an anomaly. How big a deal is that? Well, consider this analogy: In 1954, at the dawn of nuclear power, Lewis Strauss, head of the Atomic Energy Commission, promised that we were entering an age when electricity would be "too cheap to meter." Needless to say, that didn't happen, mostly because the risks of nuclear energy hugely increased its costs. But what if he'd been right? What if electricity had in fact become virtually free? The answer is that everything electricity touched—which is to say just about everything—would have been transformed. Rather than balance electricity against other energy sources, we'd use electricity for as many things as we could—we'd waste it, in fact, because it would be too cheap to worry about."

... What Mead understood is that a psychological switch should flip as things head toward zero. Even though they may never become entirely free, as the price drops there is great advantage to be had in treating them as if they were free. Not too cheap to meter, as Atomic Energy Commission chief Lewis Strauss said in a different context, but too cheap to matter. Indeed, the history of technological innovation has been marked by people spotting such price and performance trends and getting ahead of them."

My issue is that datacenter energy costs are being prioritized for commerce over residential use, so the average consumer is paying more for electricity, because a datacenter needs more electricity and they're getting tax breaks. Assuming this all improves efficiency for new products like automated robotics, there is a debatable benefit. Jevon's Paradox has no ceiling, except the environment, and people's 401ks.

meatmanek7 days ago
I just want to rant about these Artificial Analysis charts that you see everywhere:

The "most attractive quadrant" is completely meaningless. The whole point of a Pareto curve is that each point on the curve is better than everything else on at least one dimension, and that you can make these comparisons without placing a value judgement on the relative importance of the different metrics. If you make a composite score of the two metrics (any monotonically non-decreasing function, e.g. a weighted sum with non-negative weights), that score will always be maximized by one of the points on the Pareto frontier.

So going by the numbers in the 2nd chart (1st AA chart) from TFA alone:

   - there's no reason one would choose Deepseek V4 Pro 0813 (max) even though it's in the "most attractive quadrant", because GLM-5.3-Flash is both cheaper and scores better.
   - Claude Fable 5.1 (max with fallback) on the top right* could be your most attractive option if you need the best scoring model and don't care about cost, even though it isn't in the "most attractive quadrant"
   - The un-shown model off the left side of the chart could be your most attractive option if you just need lots of cheap tokens and don't care about quality.
(Obviously if you start including other factors in your score that aren't represented on the chart, then you might choose differently.)

* I also dislike the way they place the labels, and that grey line connecting the label to the point is way too subtle.

Hugsun7 days ago
Sound critique. I'll add that the Artificial Analysis intelligence index is not considered a good metric for intelligence anymore. Most of the benchmarks that it comprises are saturated or considered low signal today.
Gander57397 days ago
What is considered a good metric?
pvab37 days ago
Would it be better to have a shaded region parallel to the Pareto curve that gets darker away from it that is labeled "better value"?
Balgair7 days ago
I'm always reminded on Orwell's quote about the then new atomic bomb and his prescience on how it would all work out:

"Had the atomic bomb turned out to be something as cheap and easily manufactured as a bicycle or an alarm clock, it might well have plunged us back into barbarism, but it might, on the other hand, have meant the end of national sovereignty and of the highly-centralised police State. If, as seems to be the case, it is a rare and costly object as difficult to produce as a battleship, it is likelier to put an end to large-scale wars at the cost of prolonging indefinitely a “peace that is no peace”."

It seems, especially with open weights, that the AI is much more like the alarm clock and not the battleship. $20/mo would have been about $1 in 1944

https://www.orwellfoundation.com/the-orwell-foundation/orwel...

chr15m7 days ago
Good point. Everyone is worried about the security attack angle of AI, but it also helps the defence side. People can use it to secure their own systems.

Read the full thread on Hacker News →

Related stories