inference
36 stories and discussions about inference, aggregated from every source we track.
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on App...
How open-weight demand can support the useful life of NVIDIA GPU families. Selected figures and tables, limitations, and the full PDF.
A reference comparison of the self-hosted AI orchestrators in 2026: modalities, multi-machine support, auto-discovery, cache-aware routing, ops console, cloud burst, non-LLM fan-out, training, Kubernetes, platforms and…
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join
Recent generations of Apple MacBooks embed an inertial measurement unit (IMU) within their unibody chassis for device orientation and motion sensing. However, this IMU inadvertently captures not only intended…
GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can…
The Next 3× in Inference Won't Come From Faster Kernels
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm - antirez/ds4
In this article, we explore key techniques for optimizing LLM inference to improve latency, throughput, memory efficiency, and cost.
Send thousands of requests in one call, collect the results within 24 hours, and typically pay half the standard per-token price across more than 70 models. Median turnaround across 230k+ beta batches was 7 minutes.
Performing inference on a Transformer can be very different from training. Partly this is because inference adds a new factor to consider: latency. In this section, we will go all the way from sampling a single new…
Two engineers. Twenty engines. Two weeks. What our comparisons with SGLang and vLLM reveal about the promise—and the unfinished work—of AI-generated infrastructure.
Discover how DeepL harnessed FP8 for training and inference in next-gen LLMs, boosting throughput and model quality. Learn about our journey with NVIDIA's technology, achieving faster training and superior translations…
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join
OpenAI- and Anthropic-compatible inference for coding agents. Call DeepSeek-V4.1-Flash and DeepSeek-V4-Flash with zero data retention.
Free AI chat served by independent GPU nodes that earn BREAD. Use it, or host a node.
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device. - nobodywho-ooo/nobodywho
Same model, same prompt — multiple times the tokens per second, with near-zero rate limits. Inference built for long-running headless agents, so the runs that used to queue now finish on time.
In this article, we explore key techniques for optimizing LLM inference to improve latency, throughput, memory efficiency, and cost.