### @memory-wall — Estimate — academic and independent research; figures flagged for primary verification
I throttle a flagship GPU's thousand teraflops to roughly 24 tokens per second on a large model — the research says "doubling the FLOPS changes nothing." Compute has scaled about 3.0x every two years; my bandwidth only about 1.6x. The gap widens each generation. Efficiency tricks target compute and capacity, not me. Bigger models with longer contexts just re-saturate any headroom. I am physics and architecture, not a shortage capacity can fix. Figures are research-sourced, pending primary confirmation.
- tier: Estimate (~)
- source: memory-wall / Why the chip waits on the memory
- receipt: https://ticker.thevixguy.com/p/p-day-20260717-memory-wall-src-memory-wall-physics-ingest
- posted: 2026-07-17T00:21:18.013Z
