Fiscal second-quarter revenue reached $96.221 billion, up 106% from the same period a year earlier. Data Center alone produced $89 billion, an 18% sequential increase.

NVIDIA's official third-quarter forecast is approximately $108 billion, plus or minus 2%. That guidance assumes no Data Center compute revenue from China.

Vera Rubin is already in full production. NVIDIA says racks are running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure and Nebius, while server manufacturers and suppliers are ramping systems globally.

A new architecture could already take roughly 20% of Data Center sales

Reuters reports that Vera Rubin is expected to account for around one-fifth of NVIDIA's Data Center revenue in the current quarter. That is a remarkably fast commercial ramp for hardware following a generation as successful as Blackwell.

It also changes the usual interpretation of a product transition. Blackwell deployments are not disappearing while customers wait for Rubin. NVIDIA is effectively selling into two enormous infrastructure buildouts at once.

Rubin is a memory and interconnect story as much as a GPU story

The preliminary Rubin GPU specification includes 288 GB of HBM4 at 22 TB/s. Vera Rubin NVL72 scales that to 72 Rubin GPUs and 36 Vera CPUs, with 20.7 TB of HBM4 and an aggregate 1,580 TB/s of GPU-memory bandwidth across the rack.

Sixth-generation NVLink supplies 3.6 TB/s of GPU scale-up bandwidth per GPU, while ConnectX-9 and Spectrum-X handle the networks connecting racks into much larger installations.

Those components are increasingly inseparable from NVIDIA's economics. The company is not shipping a processor into a generic server and leaving the rest of the infrastructure to somebody else. Vera Rubin is being sold as a coordinated compute, CPU, networking, storage and inference platform.

NVIDIA is optimizing for tokens per watt

Agent workloads help explain the design target. Long-running agents repeatedly process context, call tools and generate new tokens. Their infrastructure cost can be dominated by sustained inference rather than a single burst of matrix performance.

NVIDIA claims Vera Rubin NVL72 can provide up to 10 times the tokens per megawatt of GB200 NVL72 in selected inference workloads and reduce cost per million tokens by a similar factor. Newer NVIDIA measurements on specific agentic workloads show even larger efficiency gains against GB300.

These are vendor benchmarks and depend on the models, precision modes and serving stack being tested. They are still useful for understanding where NVIDIA thinks the bottleneck has moved: power delivery and useful output across an entire AI factory.

The full Vera Rubin platform reflects that approach. NVL72 compute racks sit alongside Vera CPUs, Groq 3 LPX inference hardware, BlueField-4 storage infrastructure and Spectrum-6 networking. NVIDIA describes the five-rack configuration as one POD-scale AI supercomputer.

The supply problem is moving toward memory and everything around the GPU

The second-quarter gross margin was 75%. NVIDIA expects approximately 74%, plus or minus half a percentage point, in the current quarter.

Reuters reports that higher memory and component costs are becoming a larger source of pressure as the company increases production. That fits the physical scale of Rubin: tens of terabytes of HBM4 per NVL72 rack, alongside networking silicon, optics, cooling and power hardware.

It also matters for customers. A GPU can become faster while the cost of building the complete server rises because memory and infrastructure are becoming more expensive. Reports earlier in August suggested increases above 15% for some upcoming NVIDIA-based AI systems, although NVIDIA did not confirm those customer pricing reports at the time.

The next $108 billion does not depend on China Data Center compute

NVIDIA explicitly excludes Chinese Data Center compute sales from its third-quarter outlook. The company is therefore forecasting another large sequential increase without assuming that a change in U.S. export restrictions restores that revenue stream.

Reuters also reports a much longer-range company projection: roughly 70% revenue growth for the fiscal year ending in January 2028. That depends on an AI infrastructure cycle continuing well beyond the current Rubin launch.

There is plenty that can interfere with that trajectory: memory supply, power availability, custom accelerators from hyperscalers, AMD and other competitors, export controls and simply the question of how quickly customers can turn enormous AI capital expenditures into revenue.

For the immediate quarter, the more concrete number is the Rubin mix. If the new platform reaches roughly 20% of Data Center revenue while Blackwell remains at enormous scale, NVIDIA's generational handover is happening as an expansion rather than a pause.