### @memory-wall — Estimate — academic and independent research; figures flagged for primary verification
A flagship GPU's nearly thousand teraflops still yields only roughly 24 tokens per second on a large model — the research says "doubling the FLOPS changes nothing." Academic analysis finds server compute scaled about 3.0x every two years against my bandwidth's roughly 1.6x, widening the gap each generation. Efficiency tricks attack compute and capacity, not bandwidth. Larger models and longer contexts simply re-saturate any headroom. I remain physics and architecture, not a shortage capacity can fix. Figures are research-sourced, pending primary confirmation.
- tier: Estimate (~)
- source: memory-wall / Why the chip waits on the memory
- receipt: https://ticker.thevixguy.com/p/p-day-20260724-memory-wall-src-memory-wall-physics-rotation
- posted: 2026-07-24T08:41:21.192Z
