Are there really such architectural differences between for example sol and astra that makes astra twice more in token price? Or maybe token cost covers training expenses? For me it's hard to believe that astra need…

5 points•mtokarski•13 days ago•9 comments•
Are there really such architectural differences between for example sol and astra that makes astra twice more in token price? Or maybe token cost covers training expenses? For me it's hard to believe that astra need twice the computing power that sol needs...

9 comments

jackchillemi4 days ago
From the user side, what cost me the most wasn't the price per token, it was how many I was sending. My app first sent one giant 120k character prompt on every call, whatever the task. Now each part of the app gets its own prompt and tools, and none is over 60k characters. Most of the trimming was by hand. Claude wasn't good at it for me, it over-prompted every time.
yorwba13 days ago
Willingness to pay. If people are willing to use Astra instead of Sol even if it costs twice as much, OpenAI has no need to make it cheaper. Prices only reflect costs in a competitive market where customers actively seek out cheaper alternatives.
harsh_yadav2111 days ago
Willingness to pay is definitely the primary driver, but you aren't just paying for raw compute, you are paying for queue priority and uptime reliability under heavy concurrent load.
cyanydeez11 days ago
There still intrinsic loss to the model.
verdverm13 days ago
costs are in a large part tied to model parameter count, eg. a small model fits on one GPU, the biggest models require an entire rack

the primary difference you see through those prices is model size

that you don't seem much capability difference is why Big Ai is pushing for regulatory capture, because small models are nearly as good for many tasks at a fraction of the cost, undermining the business model they have used to justify the most expensive infra build out in human history

PaulHoule13 days ago
Also I think there are diminishing returns to larger models and the harness can make up for some weaknesses. Like maybe the huge model can hypothetically answer a question off the cuff but the small model can get some search results and think about those and give you an answer.
verdverm13 days ago
100%

I am big believer (based on personal experience) that there is a ton of ROI in harness engineering, to go along with the context engineering, invariably intertwined

kbrannigan13 days ago
It's not always larger models. Often time it's a smaller model fit for a specific use case. In the future we will see 16B models that are for : Legal, Healthcare, Science, Software Engineering, Education, Forestry, Music production

Read the full thread on Hacker News →

Related stories