🔥BTC/USDT

Nvidia previews Rubin Ultra with lower memory

Nvidia has shown major customers a revised Rubin Ultra AI accelerator configuration that preserves its projected 35-petaflop peak compute while cutting memory capacity and placing far greater emphasis on linking hundreds of chips together, according to a SemiAnalysis report circulated over the weekend.

The reported design would give Rubin Ultra 192GB of high-bandwidth memory, or HBM, compared with 288GB for the standard Rubin configuration. The reduction follows SemiAnalysis’s late-June assessment that Nvidia had dropped an earlier four-die Rubin Ultra concept in favor of a design with half as many dies.

Nvidia has not publicly confirmed the Rubin Ultra specifications described by SemiAnalysis. If accurate, the changes would show how rising HBM costs are reshaping the balance between memory capacity, networking hardware and total system pricing for the next generation of AI servers.

Memory falls as NVLink scale expands

SemiAnalysis said the previewed Rubin Ultra configuration retains the same theoretical peak performance as Rubin at 35 PFLOPs, a measurement of how many quadrillion floating-point operations a processor can perform per second. Yet the accompanying memory specification is markedly smaller.

The reported Rubin system uses 12-high HBM stacks for 288GB of memory capacity, while Rubin Ultra would use 8-high stacks for 192GB. HBM stacks place memory dies vertically, allowing AI accelerators to access data at much higher speeds than conventional server memory.

Memory bandwidth would increase by only 1 terabyte per second under the reported Ultra design, SemiAnalysis said. That offers a limited gain for workloads that move large volumes of model data continuously between memory and compute units, despite the Ultra name and substantially larger system scale.

Power requirements also remain demanding. SemiAnalysis listed Rubin Ultra’s chip-level minimum power at 1,800 watts, matching Rubin, while maximum power would rise to 2,600 watts. Those figures point to a design that seeks more aggregate performance through a larger connected system rather than a dramatic expansion in each chip’s memory resources.

The major upgrade would instead come through NVLink, Nvidia’s high-speed chip interconnect. SemiAnalysis said an NVL576 system could connect as many as 576 Rubin Ultra GPUs into one “super logical GPU,” eight times the 72-GPU world size cited for the non-Ultra configuration.

That architecture could allow customers training or running exceptionally large AI models to treat a much bigger cluster of accelerators as a unified computing resource. It also shifts more of the system’s value toward the networking equipment required to keep hundreds of GPUs exchanging data efficiently.

HBM pricing alters the system equation

SemiAnalysis attributed the redesign to the rapid increase in HBM prices and the resulting pressure on server economics. The firm estimated that HBM3 pricing rose from a low of roughly $180 to $220 per stack in the second quarter of 2025 to $600 to $700 in first-quarter 2026 contract pricing.

The report said spot prices climbed further, reaching an estimated $700 to $850 per stack in the second quarter of 2026. Such a move would sharply increase the cost of building accelerators that depend on large numbers of dense HBM stacks.

Under SemiAnalysis’s estimates, the bill of materials for a single Rubin Ultra rack initially rose from about $6.6 million to $8 million as memory prices increased. The revised specification, with lower memory capacity, could bring that cost down to about $6.4 million.

The estimated component mix changes nearly as much as the headline rack cost. SemiAnalysis calculated that HBM’s share of total system cost would decline from close to 40% to 28% after the redesign. Scale-up interconnect, meanwhile, would account for about 12% of cost, up from 4%.

The figures suggest Nvidia’s reported trade-off is designed to limit exposure to the most expensive part of the AI hardware supply chain while retaining a route to much larger model deployments. Customers would receive less memory per GPU, but could access larger pools of compute and memory across an NVLink-connected system.

Demand assumptions face a narrower test

The configuration may complicate forecasts that equate stronger AI accelerator demand with proportionately higher demand for the most advanced HBM packages. A reduction from 288GB to 192GB per accelerator would lower memory content per chip, even if Nvidia sells more networking hardware and enables larger GPU clusters.

For memory suppliers, the distinction is material. HBM demand depends not only on the number of accelerators shipped, but also on memory density, stack height, bandwidth requirements and the architecture used to connect chips within a rack or cluster.

SemiAnalysis’s cost estimates also illustrate why system designers may seek alternatives to simply adding more HBM. As memory accounts for a larger share of an AI server’s bill of materials, reducing capacity can protect rack-level pricing and leave room for spending on interconnect equipment, power delivery and cooling.

The reported Rubin Ultra changes do not indicate that demand for HBM has disappeared. They instead point to a more constrained purchasing calculation for AI infrastructure builders, where the price of memory can determine whether additional performance comes from a denser individual accelerator or from linking more accelerators together.

Nvidia’s eventual public product disclosures will determine whether the configuration previewed by SemiAnalysis reaches production in its reported form. Until then, the report has placed attention on a specific pressure point in the AI hardware market: memory pricing may increasingly dictate the design choices behind the industry’s largest computing systems.


Explore how institutional trends shape crypto liquidity in 2026—read crypto and DeFi in 2025 for macro insights.

Disclaimer: The content on this page is provided for general informational purposes only and does not represent the views or financial advice of Toobit. We make no guarantees regarding the accuracy or completeness of this information and shall not be held liable for any errors, omissions, or outcomes resulting from its use. Investing in digital assets involves risk; users should independently evaluate their financial situation and the risks involved. For further details, please consult our Terms of Service and Risk Disclosure.

Sign up and trade to earn over 15,000 USDT
Sign up