Reporting on 2B tokens of AI usage and what shapes model selection and cost

103 points•ThibWeb•about 10 hours ago•80 comments•

80 comments

epistasisabout 6 hours ago
One thing about these numbers that's absolutely shocking to me is how low the energy use is:

> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

scottcha26 minutes ago
On our journey at Neuralwatt (we are likely the provider he's using as we are the one that does all the energy observability and reporting in our cloud and I'm the CTO there) we quickly discovered that there is a large disconnect between what is in the press (like the rest of the AI narrative very subject to worst case but possibly unlikley future extrapolation of current trends) and what we see on the ground. We actually spend most of our time focused on various way energy constraints (including maximizing tokens/joule while maintaining perf) manifest in current datacenters and provide better ways to get more tokens out of the energy that are already in these data centers or already available but hard to utilize on the grid. So its really a technical constraint problem <it>today</it> rather than an impact problem and I think the press narrative could be better about this.

Regarding total costs relative to the pure energy costs it is multiple orders of magnitude different but also realize in the datacenter the energy is the pure commodity while almost every other component has huge margins driven by lack of supply. I do think over time this might get closer together (more competition on HW might lower margins) while energy might become more of a bottle neck (raising the energy prices).

gchamonliveabout 1 hour ago
> With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

This was going so well until this. Everything at scale has environmental impact because you centralize the downside and distribute the upside. This is an important alienation, but it hides the amount of heat, noise and impact on distribution that datacenters have on local infrastructure and environment.

ericdabout 2 hours ago
It's running a normal central home heat pump/air conditioner for one hour, or eight hours of playing on a gaming PC. Yeah, the propaganda around datacenter resource usage has massively outrun the truth, which makes me think there are some very interested parties pushing behind the scenes.
ThibWebabout 5 hours ago
I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
rapatel0about 1 hour ago
Yeah but performance / watt will decrease - Quadratically with improved chip scaling - orders of magnitude with improved model efficiency
CuriouslyCabout 4 hours ago
At some point the agents will be good enough that you can tell them "here's $100, go make me money" and they will, maybe not a lot and not all the time, but the EV given the cost of inference will be positive.
timmmmmmayabout 4 hours ago
"the talk of AI data centers' impact on the world" has been wildly exaggerated and you can see here that this is the least impact of any major new technology in the history of industrialization
aktenlageabout 6 hours ago
> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?

ThibWebabout 6 hours ago
Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
nxobject44 minutes ago
This wasn’t the author’s big point, but it did hit me when he said “this isn’t what we usually aspire to, but it works well for prototypes”. It captures how I feel about vibe coding - you can deal with incredibly complex tasks! But, boy, am I not about to use (say) an AI produced GPU driver on a daily basis! Perhaps we should normalize (again) “you’ll throw away your first attempt”.
disiplusabout 5 hours ago
Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading.
desmondlabout 4 hours ago
Yeah, I clicked expecting a review of GLM 5.3 Flash, but the article was more a retrospective about what he learned during his September challenge: "Only use GLM 5.3 Flash for one month"

He said his experiment was a failure because:

1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.

2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.

3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.

His takeaways from doing the experiment were:

1. Measure local usage more.

2. Experiment with agent orchestration, with bounded goals.

3. Don't count other models that are used for R&D.

4. Play with Jev.

5. Include experiments with flagship models to compare with cheap open models.

His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.

raggiabout 4 hours ago
Some vague commentary about performance with what appears to be assumptions about GPU availability, but no clarity about which inference provider is being used. If ZAI is assumed, I believe they aren't subject to the assumptions in the post based on what they've said publicly, but if they were using some other provider, perhaps.

The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.

ThibWebabout 4 hours ago
Might do more of that next time! Inference was with TensorX and Neuralwatt.
monksyabout 5 hours ago
GLM5.3-flash has been fantastic for me to make minor fixes in ambigious ways. "Fix x feature, whats going wrong. " It does the job.
surgical_fireabout 5 hours ago
GLM-5.3-flash is my implementation model after GLM-5.3 writes the plan.

It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.

monksy24 minutes ago
Same. It's barely even touching the credit I have on openrouter, it's great.
finnjohnsen2about 3 hours ago
Do you switch model mid session, or do you use subagent to do the switch after planning?

Read the full thread on Hacker News →

Related stories