Single logit inference runtime for all LLM models. Contribute to rreinold/jev-serve development by creating an account on GitHub.
2 comments
allensallinger7 days ago
Any recommendations for which of the models to try on macs with 16 GB of mem, and is this running well with any of the smaller models?
rreinold26 days ago
Guideline is (TOTAL_RAM - 5GB overhead) = max size, in billions of parameters. So 16 - 5 = 11, so <11B. Qwen 3.8 9B is a good fit
Read the full thread on Hacker News →
Related stories
- Hacker News · 2 points · 8 days ago
- Hacker News · 1 points · 9 days ago
- JEV assisted LLM Tradingdev.toDEV Community · 2 points · 10 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 7 points · 11 days ago