vllm
16 stories and discussions about vllm, aggregated from every source we track.
Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.
Building the smallest Compute Engine VM that serves Gemma 4 E2B on one Tesla T4, installing the driver and vLLM after boot, and a walkthrough of every option in the shell script that starts, checks and queries the server.
A step by step deployment of Gemma 4 E2B to a single AMD Instinct MI300X on AMD Developer Cloud, driven by Python MCP tools, and the throughput a 191.7 GiB card returns for its hourly rate.
A step by step deployment of Gemma 4 E2B with vLLM on a single Tesla T4 attached to a Compute Engine VM, and a measured comparison of the QAT w4a16 checkpoint against the bf16 reference on the same card.
How to point Oh My Pi at vLLM, llama.cpp, SGLang, and other local servers or a gateway, plus two config fixes for omp 18.2.7 and later.
How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-c
How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-c
The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.
Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex
Agentic gateway providing persistent LLM memory via a markdown wiki. Open-source core of Kortexio. - Kortexio/ContextMemory
Implement structured decision reads with DiffusionGemma on Red Hat AI to achieve low-latency, self-hosted AI for regulated industries.
The experiment Percentes runs takes one of two serving replicas away under sustained load and measures what the survivor does. Each replica is vLLM, an engin...
A technical deep dive into vLLM: PagedAttention, continuous batching, chunked prefill, CUDA graphs, and empirical benchmarks from gft-studio across GRPO rollout phases and inference serving on NVIDIA H100 and H200.