vllm

16 stories and discussions about vllm, aggregated from every source we track.

1.

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

13 points•xbill•about 11 hours ago•1 comment
2.

Building the smallest Compute Engine VM that serves Gemma 4 E2B on one Tesla T4, installing the driver and vLLM after boot, and a walkthrough of every option in the shell script that starts, checks and queries the server.

13 points•xbill•8 days ago•0 comments
3.

A step by step deployment of Gemma 4 E2B to a single AMD Instinct MI300X on AMD Developer Cloud, driven by Python MCP tools, and the throughput a 191.7 GiB card returns for its hourly rate.

11 points•xbill•13 days ago•5 comments
4.

A step by step deployment of Gemma 4 E2B with vLLM on a single Tesla T4 attached to a Compute Engine VM, and a measured comparison of the QAT w4a16 checkpoint against the bf16 reference on the same card.

7 points•xbill•12 days ago•1 comment
5.

How to point Oh My Pi at vLLM, llama.cpp, SGLang, and other local servers or a gateway, plus two config fixes for omp 18.2.7 and later.

2 points•dougcalobrisi•5 days ago•0 comments•
6.

How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-c

2 points•eatonphil•6 days ago•0 comments•
7.

How vLLM implements distribution-preserving Gumbel-max text watermarking with efficient GPU kernels, statistical detection, speculative decoding, and repeated-c

2 points•eatonphil•6 days ago•0 comments
8.

The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.

1 points•xbill•about 10 hours ago•0 comments
9.

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

1 points•xbill•about 11 hours ago•0 comments
10.

A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex

1 points•leochong•1 day ago•0 comments•
11.

Agentic gateway providing persistent LLM memory via a markdown wiki. Open-source core of Kortexio. - Kortexio/ContextMemory

1 points•vitorcastro78•2 days ago•0 comments•
12.

Implement structured decision reads with DiffusionGemma on Red Hat AI to achieve low-latency, self-hosted AI for regulated industries.

1 points•thebeardisred•2 days ago•0 comments•
13.
1 points•teleforce•3 days ago•2 comments•
14.

The experiment Percentes runs takes one of two serving replicas away under sustained load and measures what the survivor does. Each replica is vLLM, an engin...

1 points•itsveems•5 days ago•0 comments•
15.
1 points•matt_d•7 days ago•0 comments•
16.

A technical deep dive into vLLM: PagedAttention, continuous batching, chunked prefill, CUDA graphs, and empirical benchmarks from gft-studio across GRPO rollout phases and inference serving on NVIDIA H100 and H200.

1 points•gfactor_ai•9 days ago•0 comments•

Related topics