Single-forward-pass semantic routing in Rust: turn dense LLMs (Llama, Qwen, Mistral, Gemma) into a typed decision API with LoRA/QLoRA training and FP8/FP4 quantization, built on Candle. - neurono-m...

3 points•andrelgcclaudin•4 days ago•2 comments•

2 comments

waximabbax4 days ago
Its use case is still pretty narrow, just like Jev. The only time Jev could make sense is if you want thousands of requests per second and can sacrifice a bit of accuracy. Otherwise, modern LLMs are super cheap, like GPT-6 Luna, GLM-5.3-Flash, or even Gemma 4, and reason much better than Jev with better accuracy while being perfectly capable of returning structured JSON. So it’s basically enforcing typed answers with lower accuracy vs. occasionally getting the response format wrong but having better accuracy.
andrelgcclaudin4 days ago
Given the recent popularity for Jev, I also gov the simplicity behind this system, and I understood there is a big opportunity for the opensource community to create and maintain an alternative with even better powers. Then I made Typed-lm.

Typed-lm used regular LMs with Parallel Evaluation via KV-Cache Broadcasting, reaching subsecond results inference times for typed structures.

Also, it is able to train and fine tune your own models, focusing the Jev types structure, and allows providing a supply file to preffil it as a RAG.

Ideia is allowing you to create your own decision inside your servers, with privacy, and fully oriented to your business need, with Jo vendor lock-in.

Please, let me know your feedback, and add any comments and issues you'd like on GitHub.

Where is also a tutorial page:

https://neurono-ml.github.io/typed-lm/

Read the full thread on Hacker News →

Related stories