Context Language Models edit their own context, which breaks prefix caching after every edit. Suffix Cache Reuse reuses the cache of the text that survives an edit, cutting serving compute by 35% at the same accuracy.
0 comments
No comments yet.
Related stories
- DEV Community · 0 points · 13 days ago
- Covert Caches: When Is a Cache, a Cache?parallelprogrammer.substack.comHacker News · 2 points · 12 days ago
- Show HN: S3-Acceleratorgithub.comHacker News · 2 points · about 6 hours ago
- Hacker News · 1 points · 11 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 10 days ago