What an LLM server does when ten people send a prompt at once: prefill, decode, the KV cache, continuous batching, PagedAttention, and more, with animations.

3 points•mr_o47•about 10 hours ago•0 comments•

0 comments

No comments yet.

Related stories