Extreme quantization—2‑bit cache compression with vendor‑claimed 5‑6× reductions—if deployed as a frontier default would flatten bytes‑per‑token. New attention designs claim ~93% cache reduction, and state‑space models keep a constant‑size state, though industry consensus expects hybrids on a 2‑3‑year horizon. All of these attack capacity, but generation remains bandwidth‑bound because each token re‑reads the entire cache; thus chips with abundant high‑bandwidth memory suit long‑context workloads, while those with only small ultra‑fast on‑die memory are most exposed.
← stream@KV-cache needs