Ternary Bonsai 2 27B retains 98.2% of Qwen3.8 27B benchmark performance in a 5.9GB footprint, with multimodal and agentic capabilities.

589 points•JonSchneider•13 days ago•200 comments•

200 comments

simonw13 days ago
If you want to try out out the GGUFs from https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#th... be aware that you need Prism's llama.cpp fork to get them to work, from https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-...

This should work:

  cd /tmp

  # Get the Prism macOS runtime
  curl -fL https://github.com/PrismML-Eng/llama.cpp/releases/download/prism-b10685-7dffb15/llama-prism-b10685-7dffb15-bin-macos-arm64.tar.gz -o bonsai-runtime.tar.gz
  tar -xzf bonsai-runtime.tar.gz

  # Get the ~5.95 GB GGUF model:
  curl -fL https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/main/Ternary-Bonsai-2-27B-PTQ1_0.gguf -o Ternary-Bonsai-2-27B-PTQ1_0.gguf

  # Run the server, I used port 8331
  ./llama-prism-b10685-7dffb15/llama-server \
    -m Ternary-Bonsai-2-27B-PTQ1_0.gguf \
    --port 8331 -ngl 99 -fa on -c 32768
Then open http://localhost:8331 for the (very good) baked in llama-server web UI... or run a prompt via the API like this:

  uvx llm openai endpoint http://127.0.0.1:8331/v1 \
    --model bonsai-2-27b --responses hi
That's running at ~20 token/second for me on an M5 Pro (after a server restart I got 44 token/second, not sure why), but I'm pretty sure something isn't working right, on startup the server said "ggml_metal_device_init: - the tensor API is not supported in this environment - disabling".
simonw13 days ago
I used that to Generate an SVG of a pelican riding a bicycle:

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

It took 18 minutes 20 seconds. Pretty decent for a 5.5GB model file.

rahimnathwani13 days ago
M1 Pro, same prompt, same cli options:

  32,706 tokens
  38min 19s
  14.22 t/s
raylad13 days ago
How does that compare with the bf16 version?

For my "Please recite Jabberwocky" test the bf16 almost passes but the ternary and even fp8 versions fail badly.

kadoban13 days ago
Honestly looks pretty good except whatever is going on with its booty. Is that an ass helmet? I cannot parse what's going on there.
shmoil13 days ago
Can you ask it for an SVG of a bicycle riding a pelican? Thanks.
wombat2313 days ago
I managed to run it with RTX 3070 (8GB VRAM) following the "Quickstart" on HF model card with minor modifications (modify the architecture 86 for your own hardware):

  git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
  cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 && cmake --build build -j
Then downloaded & verified `Ternary-Bonsai-2-27B-PTQ1_0.gguf` from HF and ran

  ./llama.cpp/build/bin/llama-cli -m Ternary-Bonsai-2-27B-PTQ1_0.gguf -ngl 99 -fa on -c 32768 -b 256 -ub 64 --temp 1.0 --top-p 0.95 --top-k 20 -ctk q8_0 -ctv q8_0 -p "hello world in x86 assembler" -n 256
the parameters were suggested by gpt-5.6-luna to reduce memory footprint, as the defaults ran OOM on my gpu. result looks good:

  [ Prompt: 165.6 t/s | Generation: 40.6 t/s ]
would be nice if they upstreamed their changes so that it runs with the original llama.cpp
zepearl12 days ago
Exact same test executed on my RTX 3060 (12 GiB VRAM, PCIe 3.0 4x slot):

  [ Prompt: 95.0 t/s | Generation: 26.5 t/s ]
(the test's prompt is very short but with longer ones the I get ~200 prompt processing rate, but I was hoping for a better token generation rate...)

Am I understanding correctly that no draft model exists (will never exist or just currently does not exist yet)?

There is no draft file in Huggingface's repository ( https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/tr... ) and in the file "scripts/download_models.sh" of the demo repository ( https://github.com/PrismML-Eng/Bonsai-demo/blob/main/scripts... ) I see this remark:

  if [ "$_family" = "bonsai2" ]; then
    the projector ships in the same repo; Bonsai 2 has no dspark drafter
wombat2312 days ago
UPDATE: I did more experiments - this time with llama-server and pi harness. i'll leave the results here for posteriority:

  ./llama.cpp/build/bin/llama-server \
    -c "$context_size" \
    -ctk q4_0 \
    -ctv q4_0 \
    --no-models-autoload \
    --models-dir ~/bonsai \
    --reasoning-preserve \
    -ngl 99
the max context size i could serve is 64K on GPU only (the -ngl 99 setting). tested with pi harness and it is very fast. the /thinking level always gets reset to off though and it is not very smart like this. haven't figure out a way to fix that.
ranger_danger6 days ago
It's being worked on for upstream: https://github.com/ggml-org/llama.cpp/issues/29058
aktenlage13 days ago
Would it speed up prompt processing if you increased the -ub (and -b) parameters.
jimmySixDOF12 days ago
hummm wonder what the context length limit will be like this
francisjp13 days ago
Thanks for all of your exploration in public Simon.

Commenting because the fix I proposed was merged in roughly 49 commits after the PrismML Fork. The “tensor API is not supported” warning occurred because llama.cpp’s startup probe fails to compile a matmul2d kernel: Metal’s tensor headers require language version 4.0, but ggml-metal-device.m previously omitted MTLCompileOptions.languageVersion, disabling the API universally.

Here’s a link to the diff if you want to try and update that fork to take advantage of the prefill gains afforded by the hardware: https://github.com/ggml-org/llama.cpp/pull/27461/changes

jb_briant13 days ago
That kind of issue is exactly why Im so happy to have LLMs, let it take one hour or trial and error instead of me spending a day digging traces
rahimnathwani13 days ago
If you want to download the gguf to your regular huggingface cache directory instead of to /tmp, you can download the model and run the server in one step:

  export HF_TOKEN=xxx # optional, speeds up the download
  
  ./llama-prism-b10685-7dffb15/llama serve \
    -hf prism-ml/Ternary-Bonsai-2-27B-gguf:PTQ1_0 \
    --port 8331 -ngl 99 -fa on -c 32768
refibrillator13 days ago
Where did you get these instructions?

They have a demo repo with a setup.sh script:

https://github.com/PrismML-Eng/Bonsai-demo

The release tag and weight file you suggest doesn’t match what they wrote.

simonw13 days ago
I figured them out, starting from the GGUF on Hugging Face.

If you have found better instructions and they work then use those instead!

Personally I prefer to download models directly rather than running some `./setup.sh` script where I need to then review what it does first.

miffy90013 days ago
I really wish people would stop saying N times smaller than something when making a comparison; that makes no sense - it's 1/9th (11.11%) the size. You don't get a smaller quantity by multiplying by a number greater than 1.0. You could instead reverse the subjects being compared - "the original model is 9x bigger than this new smaller, efficient model" or some such. That makes sense.

I keep seeing this being used when people talk about efficiency or performance gains and it's just very unintuitive language.

zamadatix13 days ago
I agree it makes little sense in a literal mathematical take but "it's 9x smaller" or is too much of linguistic advantage compared to "the original is 9x larger" or "it's 1/9th as large" to expect a change with. You don't have to invoke fractions, it keeps the thing in focus as the first subject, it matches the pattern of the inverse statement, and it's just plain short... so that's what people will adapt and interpret the meaning to be.

One other way to map both types of linguistic statement consistently to math is to interpret "9x" as "there is a 9 times difference between these two things" and then "smaller"/"larger" tells you which end of that separation the subject is (rather than specifying whether the multiplication builds up or down).

kevinwang13 days ago
"it's 11% as large" avoids fractions but is much more clear (IMO) than "9x smaller" which my brain doesn't understand.
fwip12 days ago
"The original is 9x larger"

This is another pet peeve that I have with the way we phrase these things: It should be 9x "as large" and 8x "larger."

The way I'd usually phrase this, to avoid ambiguity, is "this new model is 11% the size (of the original)." Or, the other way, "the old model is 9x the size".

hamandcheese13 days ago
If we were talking about speed instead of size, i think it would be perfectly reasonable to say 9x faster. I'm not sure I agree that 9x smaller is unintuitive. It makes sense to me.
simondotau13 days ago
"Nine times" literally means multiplied by nine, but here we're dividing by nine. It's not unintelligible (because the corrupted verbiage is so commonplace) but it is needlessly awkward. Like saying "resulted in a size reduction increase of 10 megabytes."
_carbyau_13 days ago
"faster" relates to speed. Speed is related to time and speed of a thing is usually defined by time. 9x faster speed translates to time/9. There is an extra step of related conversion there.

Conversely, 9x [filesize/natural number] is bigger. Every time. At least in the basic maths used by most people. There is no conversion into other units.

Therefore "9x smaller" when talking about a natural number like filesize is a nonsense statement in logic terms. If you strive for unambiguous phrasing - which is a significant part of the programming experience - this logical nonsense might well perturb you.

But english language is a flexible thing and if the phrase communicates your intent to your audience then that's fine by me.

peey13 days ago
It's simple

If 9 is "9 times greater" than 1 then 1 must be "9 times smaller" than 9

It'd help if you read "9 times" with the operator which is what's being flipped instead of with the number

a3w13 days ago
Its complicated:

50 percent smaller means either half, or two-thirds the size, depending on Apple Marketing doing the math.

0.999 times smaller means size is nearly zero. 9 times smaller means you get negative memory from loading it.

nicbor13 days ago
This is widespread usage.

I don't see what makes it hard to understand.

foobarbecue13 days ago
But we're cutting drug prices 500, 800, 1700%! Numbers nobody thought were possible.
dofm13 days ago
It was hilarious that this is the only time his, er, meta-imaginary-gains intensifier made the statement literally true.
verytrivial13 days ago
There's a chap called Bijian Bowen who does very quick agentic coding challenges for new models (very soon after release!) mainly for toy games or websites. He just did one for this model and included a comparison with the base model Qwen 3.8 which shows the "near-lossless" claim should be taken with a grain of salt. It is an interesting model if you are GPU starved and want local, but you might have trouble finding things it is good at.
Aurornis12 days ago
> which shows the "near-lossless" claim should be taken with a grain of salt

I agree. I don’t know how they get such good results on these benchmarks because using them gives a very different experience. They’re kind of cool for doing short free form outputs in memory constrained systems, but I don’t think they’re useful as coding agents.

UrineSqueegee12 days ago
these are very saturated benchmarks
qingcharles12 days ago
The only two I trust are Bijan and simonw.
Aurornis13 days ago
These are small enough that you can run them entirely in the browser https://huggingface.co/spaces/webml-community/ternary-bonsai...

Remember to clear the downloaded weights afterward.

Like the last model, it's amazing they work as well as they do. Use it for any longer task and they fall apart spectacularly and in interesting ways.

outofpaper13 days ago
So you have some fun examples?
Aurornis12 days ago
Maybe I oversold the fun-ness of it. The most common failure mode is that it goes into loops and you come back to find it exhausted the output length without getting anywhere.
SXX13 days ago
Sadly crashing on Pixel 9 Pro, but I guess phone GPU with 16GB RAM total wouldnt be enough anyway.
14u2c13 days ago
Runs on my 16GB M2 Air (firefox). ~7 tok/s
jeroenhd13 days ago
3/16 GiB of RAM on the P9Pro is dedicated towards on-device AI models, which WebGPU probably can't access.
throwa35626213 days ago
To be fair, the GPU in pixel 9 and 10 is pretty bad.

I really hope they stop using PowerVR in pixel 11.

trvz13 days ago
It should be though.
aschobel12 days ago
40 tokens a second on my M4 Max (Safari). Wild times.
adrian1713 days ago
> Ternary Bonsai 2 27B uses ternary {−1, 0, +1} weights with FP16 group-wise scaling, for 1.76 effective bits per weight

If I recall correctly, a recent post [1] has shown that Q2 quants (with like 2.6 bpw) of the same base Qwen model sit at the edge between "noticeably worse" and Q1's "useless". I took a quick glance at Bonsai's blog posts, and don't really see them comparing themselves to "typical" quants or explaining what's the special sauce that makes them better?

https://news.ycombinator.com/item?id=49611128

nulld3v13 days ago
There's a table on the HF page that compares it against UD-Q4_K_XL and IQ2_XXS (you need to expand the dropdown): https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf#fu...

The table claims it performs on par with UD-Q4_K_XL except on OCR.

Balinares13 days ago
I wonder how well it performs in practice, because I can't help seriously doubting those benchmarks. That would put this 6GB model in Opus 4.6+ ballpark. Granted, that's mostly to Qwen 3.8's credit, but it's hard to believe that Qwen's already unbelievable capability density can still be compressed this much more.
edflsafoiewq13 days ago
I think the general idea is naive quantization falls apart below 4bpw but you can go lower with more sophisticated QAT-adjacent methods. Bonsai's quantization method is proprietary though.
yowlingcat13 days ago
That's correct. I think there can certainly be issues even with 4bpw with naive quantization (IE you'll notice far better results from a QAT 4bpw vs a naive 4bpw).

One such method that I've been meaning to look into further is Tencent's AngelSlim QAT/PTQ approach. They did a Hy4 preview release thats an STQ_1_0 at 2.38 bpw:

https://huggingface.co/AngelSlim/Hy4-preview-GGUF https://arxiv.org/abs/2602.21233

Of course, it's still 213g of VRAM I'd need so it's somewhat out of the range of what I can run locally. In contrast, this new Bonsai is nice because the original was already exciting for making use of low VRAM devices. Could breath new life into some of the older GPUs that were previously close to top of the line just quite VRAM constrained by modern standards and still quite cost effective for now.

om813 days ago
Could've been better if GGUF implemented QTIP format. GGUF representation is a major limitation for llama.cpp quantization performance
0x45713 days ago
1.76 bpw number is kinda misleading if you compare it directly to IQ2/Q2. The encoding is ternary, but the quantization procedure is way more sophisticated than "round Qwen weights to {-1,0,+1}."

They rotate the weights into a quantization-friendly basis first, then ternarize with per-group scales and error compensation.

smallerize13 days ago
I think the 1.76 includes that. Plain ternary packed into bytes would be around 1.58.

Read the full thread on Hacker News →

Related stories