Cross-platform pure C+GPU inference for Llama 2-family, Mistral Nemo, GPT-2, Gemma-3n, Qwen 3.x SSM, GPT-NeoX, StableLM GGUF files. Enhancements over parent repo are 99% coded by Qwen 3.6-27b with ...
PicoLM is an LLM inference engine written in C99. It currently supports llama-2, GPT-2, Qwen 3.6/3.8(+MoE) and Gemma-3n models. Significant amount of work went into CPU SIMD acceleration/testing/correctness, and wide cross-platform availability with constant testing to never lose portability (from DOS through OS/X 10.4 to modernity). CUDA/HIP is supported, and accelerated IMMA kernels are available. More work needs to be done on prompt processing speed, but text generation is quite fast already.
With v1.0-rc2, the list of tested and supported platforms became even more obscene (iPhone 1! Tru64!), and some amount of work has gone into having a Vulkan backend.
0 comments
No comments yet.
Related stories
- Hacker News · 1 points · 2 days ago
- Hacker News · 421 points · 6 days ago
- Hacker News · 3 points · 3 days ago
- Hacker News · 1 points · 4 days ago
- Hacker News · 1 points · 6 days ago
- Hacker News · 2 points · about 12 hours ago