CPU-only LLM inference engine in pure Rust — 4-bit quantized models, hand-written AVX2 kernels, speculative decoding, and a local HTTP/MCP server. No GPU, no Python runtime. - Classevelabs/rai

2 points•simonpure•about 8 hours ago•0 comments•

0 comments

No comments yet.

Related stories