Which LLM is worth it: Artificial Analysis Intelligence Index against blended API price, with the value frontier highlighted. Refreshed daily.

184 points•terryds•6 days ago•113 comments•

113 comments

ford6 days ago
You're better off going directly to artificial analysis, this is a feature-poor/misleading/outdated repackaging

Ex. this type of price estimation is quite naive - some models can require 2-3x the number of tokens to achieve the same level of intelligence. Artificial Analysis' own cost per task is a more fair estimation of cost.

pickledish6 days ago
neuronexmachina6 days ago
It's kind of crazy that claude-opus-5.5 (max) uses 119k tokens per "intelligence index task" while gpt-6-astra (max) uses 27k.

https://artificialanalysis.ai/models?cost=intelligence-vs-co...

mosselman6 days ago
Exactly. The only metric you need to look at is the cost per task and take your pick from the models that match your budget there.

Spark is such a bad model, you are far better off using Astra Medium than using Spark.

gunapologist996 days ago
Also cache price is crucial, especially for coding tasks (which, really, are where you really start to care the most about unbounded costs)
mpalczewski6 days ago
The analysis might be better on Artifical Analysis' however the way the information is presented by terrydjony is so much better. I'm able to get the exact actionable info I want. Whereas on artificial Analysis I'm analyzing and figuring it out.

It's the equivalent of ai slop writing in a data format.

Xeoncross6 days ago
If you have a 24-64GB mac, consider running Qwen3.8 27B locally at night. It's a bit slower to run locally, but if you're sleeping it's less of a problem.

Depending on your memory, you'll need to use the weaker Q4 versions but they still perform well.

It ranks higher than GPT-5.3 Codex (xhigh) or Claude Opus 4.6 (max) so is great for pairing with https://github.com/kunchenguid/gnhf for nightly experimentation, cleanup, or recommendation lists for in the morning.

vardalab6 days ago
I have all sorts of local compute, and local models fairly capable the Frontier models still way more capable/faster and local electricity consumption is something else. Good thing it is getting colder around here.

I often pair them up, and I have an Astra or Sol work as a supervisor and reviewer while Qwen 27B FP8 or Qwen 3.8 Flash Next implements things. I mostly do it as an experiment, just to see what kind of level of autonomy I can get, and they are slow to getting a decent reviewed outcome despite Qwen27B running at 100+ tps and 3.5-4K prefill rates and Qwen3.8 Next at 40 tps and 1-2K prefill. I've been also using similar approach more with OMP, not just the straight Pi harness. And OMP seems to be slower because it has more guardrails. OMP has an interesting feature where you can assign a better LLM as an advisor, wehere it just sort of monitors the progress and injects guidance. And it definitely helps, but one has to be careful. It actually turns out to be expensive if the cache reads are expensive. I learned it the hard way. Where on Fireworks' API, the cache rates for GLM 5.3 flash are quite a bit more expensive than for DeepSeek, and a simple runs ended up costing me three bucks in oversight. So a better way is to have a Frontier model running a separate tmux pane and just directing it to Wake up every 10 minutes, take a peek at what's going on, review the milestones give feedback and then sleep. This turns out to be pretty decent cost saving strategy when quote needs to be stretched. Paradoxically, OpenAI tightening up their quota allowance once they released Astra actually pushed me into all these sorts of experiments, and it's actually been interesting. I've been exploring all these smaller flash models, and it's been nice. I do like using local LLMs for chore type tasks that are just mostly information gathering, post-session reviews, stuff like that.

bitexploder6 days ago
There also exists a $600-800 GPU that can run Qwen 27B 3.8 @ like 60 t/s for around 200W of energy.

I find Qwen Flash Next quite competent as well. 27B is a solid worker like you said. If you batch work and let them crank they do remarkably well. I am working on.

I started using Herdr and taught my agents to use it. So I use OMP loop and or goal, and it has a review cycles to wake up an Opus or Sol reviewer to make sure nothing is going off the rails with a local qwen flash next coordinating for me. I kinda prefer Sol, it seems like a more patient and thorough model, especially Sol 6, but Opus 5.5 is really good and its voice and attitude is not as grating as Opus 5 for sure.

I have a few V100 GPU running Qwen 27B and they do all the work overnight. Not quite the same speeds you have yet, but this is V100 machine and an old gaming machine with 16GB 4080 and a handful of 32GB V100s... all in less than 3K (ignoring that my gaming machine is 3 years old, but runs qwen flash next for free now as I game not a lot) for my little "we have AI at home" projects and there is a lot of interest in these old GPU now because they are rolling out of data centers now.

For my local work and personal projects... they just seem to be getting done in this setup. Every few days I sit down and do a big cycle with astra/fable/opus batch things up. I have projects that are basically "i want to see what happens" to "I want this to be good, I understand the code". Some of the throwaway projects that have just kind of magically finished more or less how I wanted have been great.

tehjoker6 days ago
What do you mean by "local electricity consumption is something else"? Doesn't an M3 Pro for example draw about as much power as a bright incandescent lightbulb for a maxed out gpu workload (~100W)? That's less than a tenth of what a frontier model will use in the datacenter (which I believe are racks of BlackWell or Vera Lynn GPUs, each using 500W+).
ctkhn6 days ago
On my 64gb m3 max qwen3.8 27b has been great for planning and then letting qwen3.6 35ba3b actually implement the planned changes.
ranger_danger6 days ago
You might be interested in https://prismml.com/news/bonsai-2-27b
kolbebe6 days ago
This looked exciting until I read that it gets stuck in loops and generally wasn't a useful model
joking6 days ago
for 32gb, this is my model of reference now, you have to run it with a patched version of llama and is still not available in lmstudio or omlx, waiting for that to streamline the experience a bit. But so far, the best i had till now.
RationPhantoms6 days ago
If you're on MacOS, with atleast an M3 chip and 32GB, you should look at the splash engine.

GNHF seems exactly what I've been aiming for to handle overnight tasks.

jszymborski6 days ago
It's _so_ good, I no longer bother with Sonnet and use it locally for everything.

Consider bumping reasoning down to Medium as a default though, I agree with simonw it over thinks https://simonwillison.net/2026/Aug/16/qwen-38-27b/

felineflock6 days ago
Isn't there a way to set up a thinking budget so it automatically tells the model "Conclude your reasoning now and provide the final answer." ?
jrflo6 days ago
Does anyone actually pay API costs out of their own pocket? It's about 10x cheaper to just get a codex or chat gpt subscription, it's so heavily subsidized compared to the API that I'm sure it would be cheaper to use frontier models on a subscription plan rather than paying API prices for deepseek flash.
lowercased6 days ago
I do. For local dev work, I'm mostly using jetbrains' Junie, I can swap between a collection of models from google, openai, anthrophic.

I've had more than a few people tell me "oh, it's so much cheaper to use a $20 claude account" or "i've never hit a limit ever using my openai". Inevitably.. I end up reading/hearing "oh, I need to give it another couple hours to start using it again"... I've never hit that with my approach, even if it's costing me a bit more. Being able to work when I want when I have time has some value.

I also have openai and anthropic direct API billing set up for hosted and client projects that need to call out to an LLM service.

jrflo6 days ago
Why not use a codex or claude subscription? If you use the entire usage allotment on the $200 plan it's about $2,000 in equivalent API costs. Switching providers may be valuable but it's quite literally an order of magnitude cheaper.
WaltPurvis6 days ago
>it's costing me a bit more

Would you mind sharing how much it's costing you? I've been wanting to use APIs from within Intellij, but I hesitate because of uncertainty about the cost.

the__alchemist6 days ago
I thought Air was JB's multi-model interface? What is Junie? (I see the buttons, but am very confused by JB's AI offerings in general)

Is it worth the ~10x extra cost over the subscriptions? (This is obviously a leading question). Also, I think you can use OpenAI's subcription login with Air, but not Claude's.

handzhiev6 days ago
I do. Sharing training data with OpenAI gives me a lot of complementary tokens. I go above that but it's still quite economical and I pick the right model for the task (Luna for most).
therealdrag06 days ago
20$ subscriptions offer “top up” too I’m pretty sure.
mmmattt6 days ago
I do, 3-4$ a month of deep seek is enough for my usage
benhurmarcel6 days ago
Same thing. I'm sure a subscription would be better value if I had enough usage, but I don't. With Deepseek or similar a few $ is quite a bit of usage already.
bitexploder6 days ago
I load OpenRouter up and use models like GLM Flash 5.3, DeepSeek Flash 4.1, Luna, etc. And I often have random niche needs where I need a handful of calls for say, a really good image reader like Gemini Flash 3.8 or whatever. You can do a lot with $25 on openrouter or direct to chinese providers. I am cautious about what data I send overseas, but also like... just because it is in China does not inherently mean it is any less secure than a US provider.

I can't remember the last time any real recourse has mattered for companies getting breached or mishandling my data. Their stock just goes up and the govt just shrugs.

akmarinov6 days ago
I pay $100 for Codex and it last about a day in the weekly limit - mostly Astra and Sol.

Then i got $20 into DeepSeek and i've been using those $20 for two weeks every day now. Use case is automating computer/browser use - Astra is really good at it, but very expensive, Sol and Luna haven't been that great at it, Deepseek as at about 80% of Astra but lasts forever.

Firaxus6 days ago
Curious to hear more details about your harness and setup if you’re open to sharing.
szszrk6 days ago
I do, but via OpenRouter. Outside of work my use cases are small and cheaper models do great job at those. I noticed even if I "burn tokens like crazy" I still pay less than any subscription available (a few $ a month).

But I guess if I had an agent vibecoding on it's own, I'd go with subscription instantly.

Brendinooo6 days ago
I'd really like something that's more oriented around subscription fees.

If I want to spend $100 on LLMs next month, what should I do? Get Claude because Opus 5.5/Fable 5.1 are scoring well? Get Grok because 4.7 is supposedly a good mix of competence and cost? Try out a Chinese model? Don't do a subscription at all like this site is saying?

king_crimson6 days ago
Or get a 20$ subscription from each major provider and have 40$ spare for openrouter credits. That’s what I’m doing.
Brendinooo6 days ago
I like this idea. Is it easy to pass work around between the models?
nolok6 days ago
What actual usage you want ? Because imho the answer is very different if you want code, or computer use, or vision, or text analysis, or ...

The subscription from frontier labs can handle all that (though each has its best strenght, they all perform well on most things), but it's usually a bit of a waste compared to taking the best one for you.

danesparza6 days ago
The problem is that subscriptions with AI models get a bit 'vague' on what you get and how (or when) you might be rate-limited.
greggh6 days ago
I've been running a quant/tune of Qwen3.8 27B on my M1 Max 32gb MacBook. That plus a good pi setup is having great results. I've used a full q8 of the model before and I dont see a real difference other than how slow it is. But leaving it running overnight on tasks is working great. It is currently debugging some issues in a native Mac Swift application and getting through the list of issues just fine.

This is the one that works good for me on 32gb:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

Specifically this one: Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf

bwfan1236 days ago
I am running it on M3 pro. Works great except for prefill speed which makes it slow for many coding tasks. The newer generation of macs are promising but as a cost-sensitive user, I am also looking into cheaper 32 GB gpus from intel, AMD, nvidia. Eventually, I think these class of models will work well for most coding usecases especially given that I certainly want to have some control of the code generated.

Read the full thread on Hacker News →

Related stories