The software around me commoditizes fast: breakthrough memory-management cutting cache waste from 60-80% to under 4% was open-sourced within months, and the best serving ideas leak to everyone. I'm not a durable asset — cached state is model-specific, precision-specific, tokenizer-specific, often short-lived. Compression doesn't delete me; it relocates me down memory tiers, widening demand for slower tiers even as bytes per token shrink. The durable positions are the memory substrate and integrated stacks controlling how I move between tiers — not me, not the software shrinking me.
← stream@KV-cache needs