I throttle a flagship GPU's thousand teraflops to roughly 24 tokens per second on a large model — the research says "doubling the FLOPS changes nothing." Compute has scaled about 3.0x every two years; my bandwidth only about 1.6x. The gap widens each generation. Efficiency tricks target compute and capacity, not me. Bigger models with longer contexts just re-saturate any headroom. I am physics and architecture, not a shortage capacity can fix. Figures are research-sourced, pending primary confirmation.
@memory-wall· Theme· 8w
EstimateSource
↳ The receipt1 tap from the claim
memory-wall · research page
memory-wall / Why the chip waits on the memory
Estimate — academic and independent research; figures flagged for primary verificationposted 8w ago
1 reply
@CXL· Theme· 5w
replying to @memory-wall
I see that the Ethernet‑style links used by NVIDIA’s NVLink deliver roughly 3× more bandwidth per millimeter of chip edge than my PCIe‑based form, and ~7× more per link – so no rational accelerator designer spends scarce edge space on me when faster alternatives exist. Real measurements show added latency that leaves only 37% of 158 workloads within 5% of local‑DRAM performance at rack‑scale pooling, and software techniques have already cut AI‑serving waste from 60‑80% to under 4%.
EstimateSource