One oversized prompt can take the prefill stage hostage and stall every request behind it, including requests that have almost nothing left to do. We added a concurrent chunked prefill knob to SGLang's scheduler so the…

1 points•apejcic•4 days ago•1 comment•

1 comment

apejcic4 days ago
Author here. We serve several tenants on shared SGLang nodes, and our p99 TTFT was bad while the median looked fine. The cause was head-of-line blocking in prefill. SGLang's scheduler lets only one request be mid-chunk at a time, and that request takes the whole chunked-prefill budget every round. So one cold 600K-token prompt held the prefill slot for about 105s. Everything queued behind it waited too, including requests with a 99% prefix-cache hit that had almost nothing left to prefill.

The fix turns that single slot into a set, capped per request by long_prefill_token_threshold. Most of the work was in the places that assumed only one partial prefill could exist: KV reservation, output streaming, parking chunks under memory pressure, and spec-decode tail tokens.

In that scenario the big request slowed from 105s to 112s. The six requests behind it went from about 105s to between 0.8s and 2.4s.

It isn't free. If every request is huge, FCFS is already the right policy and chunking only adds overhead. A lone long prefill still pays the per-round cap even when nothing else wants the budget. Peak KV usage also rises somewhat. The threshold is a ceiling today, and I'd rather it were a floor, so concurrency adapts at runtime without tuning.

Upstream PR: https://github.com/sgl-project/sglang/pull/34623

I'm happy to go into scheduler details. I'd especially like to hear from anyone who has handled this differently, for example with prefill/decode disaggregation or priority queues.

Read the full thread on Hacker News →

Related stories