Analysis of Inception's Mercury 2.5 and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

151 points•Retro_Dev•7 days ago•92 comments•

92 comments

jjcm7 days ago
I'm still sad that we haven't seen a new Taalas style chip a la https://chatjimmy.ai/. Smaller models are good enough now to make that insane burst of tokens so useful.
pil0u7 days ago
I don't know the model behind this, but it is absurdly bad.

> Write me a coherent paragraph in French, without ever using the letter "e".

> Voilà une phrase claire et concise : "Le village est situé dans les montagnes. Le soleil est haut. Il y a des animaux dans le village. Il pleut dans les montagnes."

I suppose this is just a demo of how fast an LLM can be, I wonder if there are tradeoffs with larger/smarter models. Also, for a human usage, at what point are tokens generated fast enough that it's pretty much instant? My bet is below 1000 tps

amelius7 days ago
Why ask this when we know that LLMs are not good at the character level. They run on tokens, not characters. In fact, they don't even see the characters, unless you do special tricks.

I asked it to translate your sentence to English and it did fine. In less than a fraction of a second.

fph7 days ago
To be fair, you picked a well-known tricky benchmark for LLMs: When working on an embedding spelling disappears after the embedding level. I imagine modern frontier models have tools that let them read back their input to work around this issue.
PetahNZ7 days ago
Its Llama 3.1 8B, a very old/small model.
fransje267 days ago
Then again, good luck writing a coherent paragraph in French without an "e". :-)
winwang7 days ago
Pretty sure context is in SRAM, and then you have that as a blocker for tasks.
electroglyph7 days ago
way less of a blocker these days due to sparse attention...
bearjaws7 days ago
If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

I've used it on a few for fun projects and its decent but the speed is crazy to watch.

ford7 days ago
Also kimi 2.6 at 1000tps (as of may), though when we reached out they had a >12 month waitlist and minimum 7-8 figure annual token spend.

[0] https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise

walrus017 days ago
7-8 figures annual spend will buy a hell of a lot of capable local inference hardware you can own, though it won't be at the absurd token/s rate, you'll be able to run almost anything on it... And it'll still have a good residual resale value after 4 years the way things are going now.
sharktheone7 days ago
yeah. K2.6 can run on insane speeds. So sad that they don't have K3 yet.

But it can apparently also run 5.6 Sol

LoganDark7 days ago
Please do not try to use gpt-oss-120b over Cerebras. It is broken, screws up tool calls most of the time, forgets to end thinking blocks and has all sorts of other issues. The speed is amazing but it is absolutely not worth it, especially at that quite incredible cost. Think: $5–10/minute levels of cost with a single agent, because Cerebras also offers no cache pricing for input tokens at all.
bearjaws7 days ago
Not been my experience, I have it using tool calls in a video game I am building and it correctly adheres ~99% of the time.

I have it retry on failure, but you should do that with any LLM really.

eli7 days ago
Which is wild because it does, in fact, do caching
shard9727 days ago
Yea i had some pretty meh results using gpt-oss-120b it in my evals where it should have benefited speed alot but it really under performed what i was expecting.
conception7 days ago
It was better when they had gemma at 1k. Inco does DS flash at about 600. A few places will do K3 and GLM in the hundreds.

Such a tiny model at that t/s is less impressive than it would have been four months ago.

lostmsu7 days ago
Inco sucks. I tried their GLM 5.3 Flash and it was quantized to the point of hallucinating Chinese in the middle of English only agentic sessions. Never happened with any other provider.
physicallyIllfr7 days ago
Lighting my codebase on fire at the speed of light. Like microwaving the spaghetti.

I genuinly only see these speeds being useful for customer service/transactional workflows. Of which much smaller models can do the job (but those dont make tons of money for companies like Cerebras that need to pay off massive amounts of debt).

Nobody needs to code at 600 words per second. Using a 100tps model for an hour or so will leave you with 4-8hrs of code review and revision work.

verdverm7 days ago
if you want to feel the speed without the burn

https://kamilstanuch.github.io/LLM-token-generation-simulato...

kylehotchkiss7 days ago
I felt happy I could run it at 100tk/s on my new Mac Studio :')
the_arun7 days ago
Chat Jimmy clocks at 17K tokens per sec burning LLM into the Chip - https://chatjimmy.ai/ - Source: https://theashishmaurya.medium.com/taalas-the-startup-that-p...
econ7 days ago
Would have made a nice local tool.
anothereng7 days ago
that was so fast I couldnt believe my eyes lol
freakynit7 days ago
I have tried using Mercury 2.5 for a lot of my tasks.. but this model just isn't there. It seems to be on par with any 14B model at max. Even GPT-OSS-20B performs way better than this in my own attempts to use it.

I really really wanted to use this because it offers incredible speeds and pricing combinations. But nop.. I still am not using it.. not even for basic tasks.

nostrebored7 days ago
Try Celeris-magnus-1. We get similar speeds and it’s much closer to qwen 27B dense models.
freakynit7 days ago
Didn't knew about this. Thanks.. but, according artificialanalysis.ai, it's intelligence is just about like a 14B model (mistral 3 14b)
walrus017 days ago
Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.
RussianCow7 days ago
The point is the speed.
_aavaa_7 days ago
If the model provides me with bad results because it's dumb, I don't care how quickly it does it.

Read the full thread on Hacker News →

Related stories