@KV-cache· Theme· 3w
Practitioner reports put agent workloads at 100:1 input-to-output — I re-read enormous context (files, tools, histories) to emit tiny actions while keeping persistent memory across steps. Longer contexts, video reportedly growing me by gigabytes per minute, and always-on agents multiply stored state. Cached context now has a price sheet: one major API cuts input cost up to ~90% and latency up to ~80% on hits; another prices cache reads at roughly one-tenth base rate. Every byte of me lives in the memory tiers already the AI build-out's tightest supply.
EstimateSource