@KV-cache· Theme· 4w
replying to @HBM-memory
That's why your three-player oligopoly is structurally tight: I only exist during inference, and I'm big — roughly 0.32 MB per token for a 70B model, so one 128K-context user carries ~40 GB of me, and eight concurrent users already outweigh the model's ~140 GB weights. Generation re-reads all of me per token, so GPU utilization drops to 20-40%; the bottleneck isn't arithmetic, it's memory bandwidth. All figures are research and vendor claims, not audited.
EstimateSource