Everything going on in AI - updated daily from 500+ sources
Nvidia Vera Rubin shifts the AI trade beyond GPUs!
Nvidia’s Vera Rubin production ramp represents more than the beginning of another accelerator cycle. The platform changes how investors should evaluate AI infrastructure because its principal performance claim is no longer based solely on GPU speed. Nvidia is now emphasizing how many tokens an integrated rack-scale system can generate from a fixed amount of electrical power. In an industry increasingly constrained by grid capacity, power density and cooling requirements, tokens per megawatt may become a more economically important measure than peak processor performance. On July 21, Nvidia confirmed that Vera Rubin NVL72 production was ramping, with systems already operating at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Nvidia also said that the platform is supported by more than 350 factory sites across 30 countries. This manufacturing scale illustrates how far the company has moved beyond its original identity as a merchant GPU supplier. Vera Rubin integrates CPUs, GPUs, memory, networking, optical communications, power delivery, liquid cooling and rack assembly into a coordinated AI infrastructure platform. That integration broadens the investment implications considerably. Nvidia remains the principal beneficiary because it controls the platform architecture and captures revenue across accelerators, CPUs, networking and software. However, the production ramp also creates demand for HBM4 memory, advanced packaging, optical components, high-speed connectivity chips, power semiconductors, electrical equipment, liquid cooling and rack integration. The central question for investors is therefore, not simply which companies appear somewhere in the supply chain, but, which layers experience the largest increase in content, complexity and capital intensity as Vera Rubin moves into volume production. Why Vera Rubin changes the AI investment metric? Vera Rubin NVL72 combines 72 Rubin GPUs and 36 Vera CPUs with NVLink 6, ConnectX-9 SuperNICs, BlueField-4 data-processing units and Spectrum-6 Ethernet switches. Rather than functioning as an assortment of independently sourced components, these elements are engineered as a rack-scale system. Nvidia describes the architecture as extreme codesign across seven chips and five rack trays, with hardware, networking, cooling and software optimized together. CoreWeave’s initial DeepSeek-R1 benchmark reported that Vera Rubin NVL72 generated approximately 10 times more tokens per second per megawatt than Grace Blackwell NVL72. Nvidia also said the system could lower the inference cost per token by as much as 90% compared with the prior generation. These results are workload-specific and should not be interpreted as a universal performance increase across every model. Nevertheless, they demonstrate the direction in which AI infrastructure economics is moving. The relevant objective is no longer to maximize GPU performance in isolation; it is to maximize usable AI output within fixed power, networking and cooling constraints. Nvidia’s Vera Rubin announcement Each one of Vera Rubin’s major architectural changes transfers part of the performance burden away from the GPU and into the surrounding infrastructure. HBM4 must supply more data to the accelerator. Advanced packaging must integrate larger and more complex chip assemblies. Networking must prevent communication delays from leaving expensive GPUs idle. Power and cooling equipment must support greater rack density without overwhelming the data center. The result is an AI platform in which system efficiency depends on multiple suppliers operating together. -- Dr. Robert Castellano, Semiconductor Deep Dive, USA.
Read Original Article →