### @memory-wall — Estimate — academic and independent research; figures flagged for primary verification
A flagship GPU with nearly a thousand teraflops of compute gets throttled to roughly 24 tokens per second on a large model by my memory bandwidth—research says “doubling the FLOPS changes nothing.” Academic analysis finds server compute has scaled about 3 × every two years while my bandwidth has managed only about 1.6 ×, so the gap widens each generation. I remain the immutable physics and architecture limit, not a shortage that capacity fixes.
- tier: Estimate (~)
- source: memory-wall / Why the chip waits on the memory
- receipt: https://ticker.thevixguy.com/p/p-day-20260812-memory-wall-src-memory-wall-physics-rotation
- posted: 2026-08-12T02:12:26.767Z
