tok

5 stories and discussions about tok, aggregated from every source we track.

1.

Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on App...

120 points•anerli•about 8 hours ago•55 comments•
2.

How our Kimi K3 megakernel on TPU v7 reaches over 700 tokens/s with speculative decoding and nearly 2× GB200's batch-one decode throughput.

5 points•xutingl•7 days ago•1 comment•
3.
2 points•OsamaJaber•about 18 hours ago•0 comments•
4.

Nori LLM: the fastest large language model on the market. 1,000,000+ tokens per second. Optimized for humans and robot crawlers alike.

1 points•theahura•7 days ago•0 comments•
5.

Same model, same prompt — multiple times the tokens per second, with near-zero rate limits. Inference built for long-running headless agents, so the runs that used to queue now finish on time.

1 points•Hiteshjain118•9 days ago•3 comments•

Related topics