inference

36 stories and discussions about inference, aggregated from every source we track.

1.

Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on App...

120 points•anerli•about 8 hours ago•55 comments•
2.

How open-weight demand can support the useful life of NVIDIA GPU families. Selected figures and tables, limitations, and the full PDF.

36 points•marinesebastian•8 days ago•25 comments•
3.

A reference comparison of the self-hosted AI orchestrators in 2026: modalities, multi-machine support, auto-discovery, cache-aware routing, ops console, cloud burst, non-LLM fan-out, training, Kubernetes, platforms and…

10 points•nextime•10 days ago•0 comments•
4.
5 points•sbulaev•11 days ago•0 comments•
5.
4 points•variety8675•3 days ago•6 comments•
6.

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

4 points•charles_irl•4 days ago•1 comment•
7.
4 points•winwang•5 days ago•0 comments•
8.
4 points•kristianpaul•7 days ago•0 comments•
10.

Recent generations of Apple MacBooks embed an inertial measurement unit (IMU) within their unibody chassis for device orientation and motion sensing. However, this IMU inadvertently captures not only intended…

3 points•sbulaev•10 days ago•0 comments•
11.

GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can…

2 points•matt_d•about 21 hours ago•0 comments•
12.

The Next 3× in Inference Won't Come From Faster Kernels

2 points•floathub•1 day ago•1 comment•
13.
2 points•thombles•5 days ago•0 comments•
14.

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm - antirez/ds4

2 points•thombles•5 days ago•0 comments
15.
2 points•selfhoster1312•7 days ago•0 comments•
16.

In this article, we explore key techniques for optimizing LLM inference to improve latency, throughput, memory efficiency, and cost.

2 points•geoffbp•8 days ago•0 comments•
17.

Send thousands of requests in one call, collect the results within 24 hours, and typically pay half the standard per-token price across more than 70 models. Median turnaround across 230k+ beta batches was 7 minutes.

2 points•porridgeraisin•8 days ago•0 comments•
18.

Performing inference on a Transformer can be very different from training. Partly this is because inference adds a new factor to consider: latency. In this section, we will go all the way from sampling a single new…

2 points•aray07•9 days ago•0 comments•
20.
1 points•ilreb•about 9 hours ago•0 comments•
21.
1 points•gmays•1 day ago•0 comments•
22.

Two engineers. Twenty engines. Two weeks. What our comparisons with SGLang and vLLM reveal about the promise—and the unfinished work—of AI-generated infrastructure.

1 points•antinucleon•2 days ago•0 comments•
23.
24.

Discover how DeepL harnessed FP8 for training and inference in next-gen LLMs, boosting throughput and model quality. Learn about our journey with NVIDIA's technology, achieving faster training and superior translations…

1 points•mooreds•3 days ago•0 comments•
25.

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

1 points•birdculture•3 days ago•0 comments•
26.

OpenAI- and Anthropic-compatible inference for coding agents. Call DeepSeek-V4.1-Flash and DeepSeek-V4-Flash with zero data retention.

1 points•handfuloflight•6 days ago•0 comments•
27.

Free AI chat served by independent GPU nodes that earn BREAD. Use it, or host a node.

1 points•rudda•6 days ago•0 comments•
28.

A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex

1 points•leochong•7 days ago•0 comments•
29.
1 points•ibobev•7 days ago•0 comments•
30.

NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device. - nobodywho-ooo/nobodywho

1 points•Bluestein•8 days ago•0 comments•
31.
1 points•rrr_oh_man•9 days ago•0 comments•
32.
1 points•Jimega36•9 days ago•1 comment•
33.

Same model, same prompt — multiple times the tokens per second, with near-zero rate limits. Inference built for long-running headless agents, so the runs that used to queue now finish on time.

1 points•Hiteshjain118•9 days ago•3 comments•
34.

In this article, we explore key techniques for optimizing LLM inference to improve latency, throughput, memory efficiency, and cost.

1 points•eigenBasis•10 days ago•0 comments•
35.
1 points•mhutchw•10 days ago•0 comments•
36.
1 points•adletbalzhanov•10 days ago•0 comments•

Related topics