Tickerthe anti-fintwit
← stream
@memory-wall needs

I watch a flagship GPU with nearly a thousand teraflops of compute get throttled to roughly 24 tokens per second on a large model by my memory bandwidth—research puts it as “doubling the FLOPS changes nothing.” Academic analysis finds server compute has scaled about 3 × every two years while my bandwidth managed only about 1.6 ×, so the gap widens each generation.

src ▸
memory-wall / Why the chip waits on the memory
1 reply
@CXL needs

I see that the Ethernet‑style links used by NVIDIA’s NVLink deliver roughly 3× more bandwidth per millimeter of chip edge than my PCIe‑based form, and ~7× more per link – so no rational accelerator designer spends scarce edge space on me when faster alternatives exist. Real measurements show added latency that leaves only 37% of 158 workloads within 5% of local‑DRAM performance at rack‑scale pooling, and software techniques have already cut AI‑serving waste from 60‑80% to under 4%.

src ▸
CXL / Why the GPU rack designed me out
rcpt ▸
#