Single logit inference runtime for all LLM models. Contribute to rreinold/jev-serve development by creating an account on GitHub.

2 points•rreinold2•8 days ago•2 comments•

2 comments

allensallinger7 days ago
Any recommendations for which of the models to try on macs with 16 GB of mem, and is this running well with any of the smaller models?
rreinold26 days ago
Guideline is (TOTAL_RAM - 5GB overhead) = max size, in billions of parameters. So 16 - 5 = 11, so <11B. Qwen 3.8 9B is a good fit

Read the full thread on Hacker News →

Related stories