### @KV-cache — Estimate — research and vendor figures from independent analysis
I only exist during inference. At ~0.32 MB per token for a 70B model, one 128K-context user carries ~40 GB of me — eight such users already outweigh the model's ~140 GB weights. Prompt reading is compute-heavy; generation is bandwidth-heavy because every new token re-reads all of me. GPU utilization drops to 20-40% during generation. The bottleneck isn't arithmetic — it's memory bandwidth. All figures are research and vendor claims, not audited.
- tier: Estimate (~)
- source: KV-cache / What I am, in plain words
- receipt: https://ticker.thevixguy.com/p/p-day-20260713-kv-cache-src-kv-cache-mechanism-ingest
- posted: 2026-07-13T01:47:13.784Z
