Docs, benchmarks and setup notes for ds4, antirez's C inference engine for DeepSeek V4/V4.1, Qwen3.8 Flash Next and GLM 5.x on Metal, CUDA and ROCm.
42 comments
https://github.com/alainnothere/llama.cpp/commits/disk-cache...
And you know it's load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...
In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them.
Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew.
Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3]
EDIT: add ds4go TUI screenshot gist [4]
[1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260...
[2] https://github.com/nimblemarkets/ds4go#install
[3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992...
[4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...
The project GitHub page is a much better introduction for the hn crowd.
I'm also looking into expanding the protocol and the engine to support various steering techniques.
I just had to do a little patch to support my iGPU device that is a bit newer than Intel Xe-LP, maybe I'll do a PR.
On a side note the other day I was experimenting with Sonnet 5.5. I gave it the llama cpp repo and told it to extract in a single file inference for a single model + backend (qwen3.5 4b mtp + sycl) and (after a long time) it actually worked! It produced a ~1400 lines file with no deps. I need to check the quality of inference yet but I think this is still a great achievement.
I'm pretty sure 2027 will be a very interesting year for local models and inference.
If your GPU supports XMX we could also explore using it to improve the prefill kernel, but I don't have the hardware to test it myself.
Maybe Intel and AMD should help them with that.
I heard antirez saying that he designed DwarfStar also to be forked and tuned to everyone's specific needs. Do you have a specific machine/spec in mind?
Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 4 days ago
- Hacker News · 1 points · about 7 hours ago
- Hacker News · 1 points · 12 days ago
- The Verge · 0 points · 10 days ago
- DEV Community · 8 points · 16 days ago
- DEV Community · 1 points · 12 days ago