Same model, same prompt — multiple times the tokens per second, with near-zero rate limits. Inference built for long-running headless agents, so the runs that used to queue now finish on time.

3 points•Hiteshjain118•9 days ago•5 comments•

5 comments

Hiteshjain1189 days ago
We built a custom inference engine to deliver fast and cheap inference. We optimized this engine for coding and research workload and achieving an average speed of 469 tok/s and 4.5x lower costs than Openrouter. Software development at that speed feels different. Get an API key and try in your Opencode!
garynamtu9 days ago
Great service and it runs quite fast !!
digestainews9 days ago
What are the limits of the free version?
Hiteshjain1189 days ago
We offer $5 free credits. But here's a $20 code for sign up from hackernews: HACKERNEWS20
ryanjosebrosas9 days ago
We tested it earlier! Speed and cache are all good!

Read the full thread on Hacker News →

Related stories