A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new…

1 points•theanonymousone•1 day ago•0 comments•

0 comments

No comments yet.

Related stories