1 comment
gauravapisceanabout 3 hours ago
Author here. We tested a simple controller designed to reduce pauses while an LLM generates a response.
In one setup, it reduced P99 inter-token latency by 27.7%. But it did not work on larger models or multiple GPUs. We traced the failure to the timing signal used by the controller.
We published both the positive and negative results because the failure identifies an important limitation and suggests what a better controller should measure.
Happy to answer questions about the implementation, experiments, or results.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 9 days ago
- InputConfig: Free Controller Mapper for Macinputconfig.comHacker News · 1 points · 3 days ago
- The Verge · 0 points · 10 days ago
- Hacker News · 25 points · 10 days ago
- Hacker News · 3 points · 7 days ago
- Hacker News · 2 points · 1 day ago