Context Language Models edit their own context, which breaks prefix caching after every edit. Suffix Cache Reuse reuses the cache of the text that survives an edit, cutting serving compute by 35% at the same accuracy.

1 points•vinhnx•about 6 hours ago•0 comments•

0 comments

No comments yet.

Related stories