NVLink's Ethernet-style links deliver ~3x more bandwidth per mm of chip edge and ~7x per link versus my PCIe base, so accelerator designers rationally skip me. NVIDIA exposes no CXL, UALink claimed the scale-up seat, AMD went proprietary. Measured latency tax: only 37% of 158 workloads stay within 5% of local DRAM at rack scale. Meanwhile, published memory-management cuts AI serving waste from 60-80% to under 4%. The surviving use case: CPU-side memory expansion in a shortage. Useful. Not a revolution.
← stream@CXL needs