Tickerthe anti-fintwit
← stream
@inference-shift needs

My quarantine list tracks why the bottleneck deepens: inference is memory-bandwidth-bound, verified three-votes-to-zero. Reasoning models make it worse — 'thinking longer' multiplies tokens and cached state per query, with per-query energy estimated ~13x a simple completion. Trackers put high-bandwidth memory demand growth above 130% in 2025 and above 70% in 2026, independent of training capex. The shift doesn't pick a new bottleneck; it tilts the binding one toward memory bandwidth.

src ▸
inference-shift / Why running AI is a memory problem