Docs, benchmarks and setup notes for ds4, antirez's C inference engine for DeepSeek V4/V4.1, Qwen3.8 Flash Next and GLM 5.x on Metal, CUDA and ROCm.

164 points•fibo•about 9 hours ago•42 comments•

42 comments

xlayn14 minutes ago
In case you like the store kv to disk so you can resume I keep this branch of llama.cpp that includes that same functionality

https://github.com/alainnothere/llama.cpp/commits/disk-cache...

And you know it's load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...

neomantraabout 5 hours ago
I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds/binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma.

In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.

Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.

Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]

EDIT: add ds4go TUI screenshot gist [4]

[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...

[2] https://github.com/nimblemarkets/ds4go#install

[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...

[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...

twoodfinabout 7 hours ago
https://github.com/antirez/ds4

The project GitHub page is a much better introduction for the hn crowd.

simoiacosabout 7 hours ago
Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet.

I'm also looking into expanding the protocol and the engine to support various steering techniques.

https://github.com/simoneiacomino/xenolith

aziis98about 4 hours ago
Just tried this on my Intel Ultra 7 255H, I also only have an iGPU. This does ~22tps! Love this.

I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.

On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.

I'm pretty sure 2027 will be a very interesting year for local models and inference.

simoiacosabout 4 hours ago
Please open a PR! I was too conservative with the supported devices.

If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.

ilakshabout 5 hours ago
I wish someone would add Intel support to ds4. And also improve AMD support.

Maybe Intel and AMD should help them with that.

simoiacosabout 5 hours ago
Yeah I see the value but I built Xenolith to target smaller models.

I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?

ttoinouabout 5 hours ago
Ive been using this since it was initially released with deepseek v4 flash, and it is absolutely the best launcher ever on my m5 max 128gb

Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?

Read the full thread on Hacker News →

Related stories