One oversized prompt can take the prefill stage hostage and stall every request behind it, including requests that have almost nothing left to do. We added a concurrent chunked prefill knob to SGLang's scheduler so the…
1 comment
The fix turns that single slot into a set, capped per request by long_prefill_token_threshold. Most of the work was in the places that assumed only one partial prefill could exist: KV reservation, output streaming, parking chunks under memory pressure, and spec-decode tail tokens.
In that scenario the big request slowed from 105s to 112s. The six requests behind it went from about 105s to between 0.8s and 2.4s.
It isn't free. If every request is huge, FCFS is already the right policy and chunking only adds overhead. A lone long prefill still pays the per-round cap even when nothing else wants the budget. Peak KV usage also rises somewhat. The threshold is a ceiling today, and I'd rather it were a floor, so concurrency adapts at runtime without tuning.
Upstream PR: https://github.com/sgl-project/sglang/pull/34623
I'm happy to go into scheduler details. I'd especially like to hear from anyone who has handled this differently, for example with prefill/decode disaggregation or priority queues.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 1 day ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 2 points · 8 days ago
- Multi-Agentdevelopers.openai.comHacker News · 2 points · 1 day ago
- DEV Community · 5 points · 12 days ago
- DEV Community · 11 points · 8 days ago