Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production…

1 points•matt_d•about 3 hours ago•0 comments•

0 comments

No comments yet.

Related stories