For the past few days, I’ve been testing the (currently) top-of-the-line M5 Ultra Mac Studio with 256 GB of RAM. I’ll cut to the chase: the M5 Ultra Mac Studio is a dream machine for local AI agents. This computer…

269 points•piotrgrabowski•10 days ago•262 comments•

262 comments

simonw10 days ago
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:

  Qwen3.8 27B tokens/sec generation speed

  Prompt size    8K    64K   128K   256K
  RTX 5090 PC    59    51    44     n/a
  M5 Ultra       48    39    32     24
  M3 Ultra       31    23.5  20     15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
gpugreg10 days ago
Those RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
beastman829 days ago
can confirm.

I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.

liuliu9 days ago
Both are probably single-token decode performance, which is reasonable to show. Otherwise agree RTX 5090 should shinebetter with NVFP4.
GeekyBear9 days ago
The issue is that the moment you want to run the more capable models that will no longer fit in a single 5090's memory, performance falls off a cliff.
searealist9 days ago
... or with llama.cpp with MTP.
peri-cl10 days ago
Those are some incredible graphs, that leap in prompt processing going from M3 to M5.

Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think).

/meta Here's a CSS filter that stops those nuisance chart animations,

    macstories.net##*:style(animation: none !important; transition: none !important)
redox9910 days ago
A dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
tcdent9 days ago
A dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform.

Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.

peri-cl10 days ago
They do MoE. They benchmarked GLM 5.3-flash (320B / 18B), and Qwen 3.8-flash-next (125B / 6B). The dense Qwen is only focused (I assume) because it's about the only thing that fits on a 5090, that they can compare the two heads on.
skohan9 days ago
1.2 T/s is not that modest is it? That's very close to an RTX pro 5000
nacs9 days ago
That's a dense model. Of course it will do worse.

Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).

peri-cl9 days ago
Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,

https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...

(Note it's a sparse MoE with only 6B active).

karmakaze9 days ago
I really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in.

These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.

srcreigh10 days ago
This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware.

It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization.

The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents.

An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.

zozbot2349 days ago
Astra-Ultra? Even the largest open model to date (Kimi K3) is nowhere close to Astra level, and it will be quite slow even on the highest-spec M5 Ultra, with achievable speeds of about 0.5 tok/s at most due to having to stream weights from SSD (~13 GB/s on the highest storage capacity M5 Max machines so far). This is OK for doing simple Q&A in the background but it's far from a genuine coding experience. You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights (and this is where the "Ultra" part sort of becomes relevant; Kimi series models have good support for agent swarms) but this would decrease single-session performance even further. It would only be usable for background jobs, though the hardware would then have a chance of paying for itself if it was fully used on a 24/7 basis.
srcreigh9 days ago
> You'd have to test batching of multiple thinking streams in order to try and raise overall tok/s via layer-wise reuse of the streamed weights

isn’t this very straightforward to do..? I thought batching for Qwen models is already proven out.

> but this would decrease single-session performance even further

Well let’s take Qwen 3.8 27B. Throughput for M3 at 8 agents is 4x compared to single agent. [1]

It’s really not clear to me that 8 concurrent agents at half speed will be worse task completion latency than 1 agent.

And that’s M3 studio benchmarks, not even M5 ultra, and without the many software improvements we will see

If you haven’t tried Qwen 3.8 27B xhigh on a task you might not get the hype. Idk.

If you’ve tried doing this and don’t like it sure, and be specific about what isn’t effective, but let’s not speculate.

[1]: https://omlx.ai/benchmarks/performance/69kzkrv8?utm_source=c...

slowin9 days ago
> This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan.

Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model.

That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).

nowittyusername9 days ago
With the latest codex (weekly quota burn) fiasco I tried open weight alternatives for the first time. And tyeah... open weight models cant compete with likes of astra yet. But, my hope is that by the time I get my Mac studio at end of november an open weight models would have closed the gap (which i think is realistic at the speed of progress). Now its true a better gpt version will also be available then but it also seems the gap is shrinking with time so theres that.
_hugerobots_9 days ago
Local models can be widely used as productive assets. Yes the infrastructure of SOTA API models is engineered specifically for you to be that utility, but the blanket statement that local isn't up to par is intensely short sighted. Billions of tokens per month on local pays for the hardware when compared to sota costs per month.
sajithdilshan10 days ago
On Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$.

That’s like 12 years worth of OpenAI Pro subscriptions

11223310 days ago
Hard to guess, it can go either way. If you will need to be in a syndicate to use non-sterilized models, that mac makes sense. But if there is mandatory registration of personal cyberarms, you risk going to mines once they check you purchases. You could try to play normie and pretend you simply wanted to show off, by keeping your actual work on external disk, but that leaves traces on system. Counting on someone in the Gap renting you gray iron works as long as you can swap credits. Still, this gear is tiny. Put it in your e-car, with uplink, and leave it at uncle's farm. Discreet.
woah9 days ago
It was a dark rainy night in Neo-Tokyo as Blake puffed on his vapor cartridge and watched the Mac dealers prowl below. Almost 15k Union Credits to get one of them to meet you in an e-cafe with a fully loaded M5, but man, the inference rush from one of those things was something else.
Razengan10 days ago
I gotta have some of what you had :)
glitchc9 days ago
I'm sold on "personal cyberarms" as a concept

Do they include footguns from pointer bugs?

nowittyusername9 days ago
512 option isnt worth it imo, you get severe slowdowns when weights are that large. 256 is the sweet spot, you can run large open weight models at decent speeds for full private inference.
cma9 days ago
> you get severe slowdowns when weights are that large.

Not necessarily for MoE

Octoth0rpe9 days ago
a) we don't actually know what the prices will look like yet, b) what about same weights + huge context? or, same weights that you'd run on 128gb/256gb, but multiple models running for different tasks?
throw0101c9 days ago
> 512 option isnt worth it imo, you get severe slowdowns when weights are that large.

I think most people are getting 512 for running Chrome with a bunch of tabs open. /s

geodel10 days ago
Agreed.

Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.

Kurtz7910 days ago
I think we all expect the heavy subsidized subscriptions to end or significantly increase in price at some point, but it could be years from now and I'd rather spend a similar figure on an hypotetical Mac Studio M8 Ultra, or whatever more advanced competitor that will have likley appeared by that time.

A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.

vardump10 days ago
I hope that was sarcasm.
jeffybefffy5198 days ago
And not have to worry about them training on your data even tho you ticked a box saying "dont do that".
patrickmcnamara9 days ago
HN always has these completely contrived counterarguments. What is actually going to realistically happen that will prevent use of an LLM provider? Did you think that the OP literally meant the 12 years or maybe it was just to show how expensive using a Mac Mini as an alternative is?
ericmay9 days ago
Just commenting here because you're discussing hardware: I thought the test results from the SSD published in this article [1] were pretty interesting. Maybe that's old news though.

[1] https://www.macworld.com/article/3238319/mac-studio-m5-max-r...

simonw10 days ago
Yeah, anyone who thinks local AI is going to save them money is likely to be disappointed, at least if they want to run models that are even remotely capable.

Plenty of other reasons to get excited about local AI, but I don't think cost is one of them.

criddell10 days ago
Maybe you are using a local model to go after some Millennium Prize problem and you don't want OpenAI to take your work and use it to win the prize for themselves? $15k might be a bargain.

And, yes, I know a current local model wasn't going to solve the Navier-Stokes problem, but I'm just using it as an example where privacy might be valuable.

hgoel9 days ago
Despite being on a site called Hacker News, we seem to often overlook the simple aspect of wanting local AI hardware to hack (not necessarily in the cybersecurity sense) with. I got my local AI hardware because it's an enjoyable hobby for me.
matt-p9 days ago
On a personal level maybe not yet, but for a medium business upwards it may make sense.
ApolloFortyNine10 days ago
The model being tested is 18k as configured.

I didn't expect this to make the 5090 to look like a good deal.

nacs9 days ago
5090 has 32GB VRAM.

It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.

orsorna9 days ago
Is it that silly? You could run multiple 27B models in parallel.
asimovDev9 days ago
can run multiple subagents of Qwen 27B though, right? Unless I am fundamentally misunderstanding how VRAM constraints work
tempoponet10 days ago
While I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks.

This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.

_hugerobots_9 days ago
Speed vs task-completion is a new conversation and a great point. Whereas the cost to compute doesn't exist in a vacuum, making mistakes costs less, is easier to maintain with granularity and a whole host of other factors when you own the lab.

Read the full thread on Hacker News →

Related stories