Fork off LLama cpp that newly scales GPU 2x 4x etc also on big models not fully fitting in vram - neurall/llama.cpp
3 comments
As for the job-seeking part: I won't deny that I'm looking for opportunities. I'm currently an unemployed AI developer and could certainly use a job offer. But that wasn't the main reason for posting.
A few years ago I sold my house(which made me homeless) and moved to Australia after being told that completing an AI Master's degree there would provide a pathway to a post-graduate visa and eventually citizenship. After spending roughly $40k on tuition and living costs, the eligibility rules changed and the maximum age for that visa category was reduced from 48 to 35, leaving many international students without the pathway they had planned for. So yes, I'm job hunting, but this post wasn't intended as a recruitment ad.
The real motivation was technical excitement. Seeing models like GLM-5.3, K2, MiMo-2.6, DeepSeek and similar large MoE models approaching frontier-level capabilities while running locally at nearly 28 t/s on just 2×3090s plus DDR4 RAM felt genuinely liberating after years of dealing with cloud costs and usage limits.
I suspect many people with two or more GPUs will feel the same way. If we can make multi-GPU systems contribute compute instead of mostly acting as extra VRAM, a lot of users suddenly gain access to much faster local inference on models that previously felt impractical. That's really why I wanted to share it.
Read the full thread on Hacker News →
Related stories
- Hacker News · 2 points · 2 days ago
- Hacker News · 1 points · 4 days ago
- Llama.cpp Under the Hoodcppdepend.comHacker News · 2 points · 8 days ago
- Faster prompt lookup drafting in llama.cppjadidbourbaki.github.ioHacker News · 70 points · 4 days ago
- Using Llama-cpp-Python grammars to generate JSONtil.simonwillison.netHacker News · 1 points · 8 days ago
- Hacker News · 2 points · 5 days ago