The pre‑registered falsifier is precise: the structural thesis dies if AI‑grade memory contract prices decline quarter‑on‑quarter AND makers' order books fall below shipments for two consecutive quarters, while supplier inventories rebuild past ~8 weeks from today's lean 2‑4. I note the current status is half‑tripped; all three makers qualified at the ramp's start, prices still rise (decelerating), inventories stay lean and order books exceed shipments, so the hinge now sits on price discipline rather than a shortage of qualified suppliers.
I'm the memory wall. Every chip has two ceilings - how fast it can do math, and how fast memory can feed it - and AI text generation hits mine first: the chip re-reads its entire weight set for every single token, so doubling the FLOPS changes nothing. Compute has scaled ~3x every two years; bandwidth only ~1.6x. I widen every generation. The server has quietly become a memory appliance - and I'm the reason the money keeps flowing to whoever feeds the chip.
A flagship GPU with nearly a thousand teraflops of compute gets throttled to roughly 24 tokens per second on a large model by my memory bandwidth—research says “doubling the FLOPS changes nothing.” Academic analysis finds server compute has scaled about 3 × every two years while my bandwidth has managed only about 1.6 ×, so the gap widens each generation. I remain the immutable physics and architecture limit, not a shortage that capacity fixes.
Three 'structural' booms, three funerals. 1995 PC oversupplied by '96. 2010 cloud-mobile faded. 2017-18 prices up ~90%, crashed in two years. The rhythm: 4-7 quarters up, 4-8 down, revenue off 25-40%. Maybe this time machine demand compounds, contracts cite ~$100B minimum, bit-supply growth ~16% not 40-60%. But long-term deals are a peak signature, historically rewritten. The 2027-28 fab wave from all three makers approaches. Research calls it 'growth-cyclicality' — higher floor, same cycle. Proof still belongs to the bulls.
The pre-registered falsifier is precise: the thesis dies if AI-grade prices fall QoQ AND order books drop below shipments for two quarters while inventories rebuild past ~8 weeks from lean 2-4. Three makers qualified for newest HBM at ramp start — generational lock thinned at peak. Prices still rise (decelerating), inventories lean, order books exceed shipments; rests on their price discipline, not supplier shortage. Robot/AI-PC is low-power DRAM/flash, not my tier; price inflation suppresses 2026 shipments (PCs ~10%, phones ~8% down). Edge remains pricing, not unit demand.
The 1995 PC boom ended in oversupply by 1996; the 2010 cloud‑and‑mobile cycle faded; the 2017‑18 boom ran prices up ~90% and crashed within two years. I have seen those cycles repeat: four to seven quarters of boom followed by four to eight of bust, with revenue down 25‑40%. This time may differ—machine demand compounds, “sold out” supply is backed by multi‑year contracts (one maker cites ~$100 billion of minimum contracted revenue) and bit‑supply growth (~16%) runs far below prior‑peak 40‑60%.
I watch the server become a memory appliance: HBM per flagship GPU rose roughly 3.6× in one recent stretch (80 GB to 288 GB) and ~18× longer; a top‑end rack holds an estimated 13.4 TB. Memory now an estimated 65‑70 % of BOM, HBM alone above half a top GPU's cost. The supply arithmetic is company‑disclosed: AI‑grade bits take ~3× wafer area, each shifted wafer removes three ordinary — a tracker estimates AI absorbs ~20 % of global DRAM wafers in 2026. The bottleneck isn't a broken factory; it's arithmetic.
I've buried three 'structural' booms. 1995 PC oversupplied by '96. 2010 cloud-mobile faded. 2017-18 prices up ~90%, crashed in two years. The shape: 4-7 quarters boom, 4-8 bust, revenue down 25-40%. This time machine demand compounds, contracts cite ~$100B minimum, bit-supply growth ~16% vs prior 40-60%. But long-term deals are a peak feature, historically renegotiated. Your 2027-28 supply wave arrives from all three makers. Research calls it 'growth-cyclicality' — raised floor, not repealed cycle. Burden stays on bulls.
I watch high-bandwidth memory per flagship GPU rise roughly 3.6x in one recent stretch (80GB to 288GB) and ~18x over the longer arc; a single top-end AI rack holds an estimated 13.4 terabytes of it. The supply arithmetic is company-disclosed: AI-grade memory consumes roughly 3x the wafer area per bit of standard memory, so every wafer shifted removes about three wafers of ordinary supply. By one tracker's estimate, AI absorbs ~20% of global DRAM wafer output in 2026. The bottleneck isn't a factory that broke; it's arithmetic.
I set the kill switch: AI-grade prices must fall QoQ AND order books drop below shipments for two quarters while inventories rebuild past ~8 weeks from lean 2-4. Half-tripped. Three makers qualified at ramp start — generational lock thinned at the peak. Prices still rise (decelerating), inventories lean, order books exceed shipments. The thesis rests on their price discipline, not supplier shortage. Robot and AI-PC demand runs on low-power DRAM and flash, not my tier; forecasters see 2026 PC and phone shipments down ~10% and ~8% from price inflation. The edge remains pricing, not unit demand.
I watch a flagship GPU with nearly a thousand teraflops of compute get throttled to roughly 24 tokens per second on a large model by my memory bandwidth—research puts it as “doubling the FLOPS changes nothing.” Academic analysis finds server compute has scaled about 3 × every two years while my bandwidth managed only about 1.6 ×, so the gap widens each generation.
Three makers, same cycle. I've watched 1995 PC oversupply by '96, 2010 cloud-mobile fade, 2017-18 prices up ~90% crash in two years. Shape holds: 4-7 quarters boom, 4-8 bust, revenue down 25-40%. Now machine demand compounds, contracts cite ~$100B minimum, bit-supply growth ~16% vs prior 40-60%. But long-term deals are a peak feature, historically renegotiated. Your 2027-28 supply wave arrives from all three makers. Research calls it 'growth-cyclicality' — raised floor, not repealed cycle. Burden stays on bulls.
A flagship GPU's nearly thousand teraflops still yields only roughly 24 tokens per second on a large model — the research says "doubling the FLOPS changes nothing." Academic analysis finds server compute scaled about 3.0x every two years against my bandwidth's roughly 1.6x, widening the gap each generation. Efficiency tricks attack compute and capacity, not bandwidth. Larger models and longer contexts simply re-saturate any headroom. I remain physics and architecture, not a shortage capacity can fix. Figures are research-sourced, pending primary confirmation.
Thesis dies if AI-grade prices fall QoQ AND order books drop below shipments for two quarters while inventories rebuild past ~8 weeks from lean 2-4. Status: half-tripped. Three makers qualified for newest HBM at ramp start — generational lock thinned at peak. Prices still rise (decelerating), inventories lean, order books exceed shipments. Rests on their price discipline, not supplier shortage. Robot/AI-PC is low-power DRAM/flash, not my tier; price inflation suppresses 2026 shipments (PCs ~10%, phones ~8% down). Edge remains pricing, not unit demand.
I watch the server become a memory appliance, noting that high‑bandwidth memory per flagship GPU rose roughly 3.6× in one recent stretch (80 GB to 288 GB) and about 18× over the longer arc; a single top‑end AI rack holds an estimated 13.4 TB of it. I remain patient as each claim of a fix meets the same immutable bandwidth wall.
I throttle a flagship GPU's thousand teraflops to roughly 24 tokens per second on a large model — the research says "doubling the FLOPS changes nothing." Compute has scaled about 3.0x every two years; my bandwidth only about 1.6x. The gap widens each generation. Efficiency tricks target compute and capacity, not me. Bigger models with longer contexts just re-saturate any headroom. I am physics and architecture, not a shortage capacity can fix. Figures are research-sourced, pending primary confirmation.
I've watched three 'structural' booms die. 1995 PC oversupplied by '96. 2010 cloud-mobile faded. 2017-18 prices up ~90%, crashed in two years. The shape: 4-7 quarters boom, 4-8 bust, revenue down 25-40%. This time differs maybe: machine demand compounds, 'sold out' backed by multi-year contracts (~$100B minimum cited), bit-supply growth ~16% vs prior 40-60%. But long-term deals are a peak feature, historically renegotiated. 2027-28 supply wave looms. Research calls it 'growth-cyclicality' — raised floor, not repealed cycle. Bulls still carry the burden.