Fork off LLama cpp that newly scales GPU 2x 4x etc also on big models not fully fitting in vram - neurall/llama.cpp

1 points•neuralll•5 days ago•3 comments•

3 comments

neuralll5 days ago
I built a llama.cpp fork that fixes a real problem: with stock llama.cpp, extra GPUs mostly just add memory, not speed, for MoE models too big for VRAM only one GPU computes at a time. Mine adds a live expert cache so multiple GPUs actually compute in parallel with the CPU. Measured on 2x RTX 3090: 12.3 to 27.7 t/s on GLM-5.3-Flash (117GB), 4.7 to 9.9 t/s on MiMo-2.6-Flash (132GB) — 2-2.3x, same output quality. Builds on an existing open PR (csantiago78's GPU expert-cache PR #27861) — credited prominently in the README, along with GLM-5.3-Flash support from timkhronos's PRs. Happy to answer questions about the implementation. Looking for Job abroad too as war is looming in EU
fshr5 days ago
Are you going to / have you tried to contribute this to the main repo? Is the fork for visibility for job hunting?
neuralll5 days ago
I'm definitely trying. I've already posted the work in the llama.cpp PR thread, and a few collaborators have started testing it. Hopefully we can turn it into a proper PR and get it merged upstream.

As for the job-seeking part: I won't deny that I'm looking for opportunities. I'm currently an unemployed AI developer and could certainly use a job offer. But that wasn't the main reason for posting.

A few years ago I sold my house(which made me homeless) and moved to Australia after being told that completing an AI Master's degree there would provide a pathway to a post-graduate visa and eventually citizenship. After spending roughly $40k on tuition and living costs, the eligibility rules changed and the maximum age for that visa category was reduced from 48 to 35, leaving many international students without the pathway they had planned for. So yes, I'm job hunting, but this post wasn't intended as a recruitment ad.

The real motivation was technical excitement. Seeing models like GLM-5.3, K2, MiMo-2.6, DeepSeek and similar large MoE models approaching frontier-level capabilities while running locally at nearly 28 t/s on just 2×3090s plus DDR4 RAM felt genuinely liberating after years of dealing with cloud costs and usage limits.

I suspect many people with two or more GPUs will feel the same way. If we can make multi-GPU systems contribute compute instead of mostly acting as extra VRAM, a lot of users suddenly gain access to much faster local inference on models that previously felt impractical. That's really why I wanted to share it.

Read the full thread on Hacker News →

Related stories