AI inference is turning data-center storage into a more central constraint on system design, pushing demand beyond GPUs and high-bandwidth memory toward high-capacity NAND flash that can hold far larger pools of frequently accessed data. Sandisk is positioning its proposed High Bandwidth Flash architecture as one response, arguing that NAND-based storage can sit alongside HBM in inference servers where capacity, cost per bit, and data proximity matter as much as raw compute throughput.
The shift reflects how AI workloads change once models move from training environments into production. A system handling search queries, enterprise agents, multimodal prompts, or real-time recommendations must repeatedly retrieve model parameters, embeddings, cached outputs, user context, and source data. Those requirements create a storage hierarchy: HBM supplies extreme bandwidth near accelerators, DRAM supports active working memory, and NAND provides a much larger, lower-cost layer for data that must remain readily available.
Sandisk said at its 2026 investor event that enterprise data-center flash demand could reach 1.2 zettabytes by 2030. The company projected mid-to-high double-digit annual revenue growth between fiscal 2028 and fiscal 2030, extending a strategy that increasingly ties its data-center business to AI infrastructure rather than consumer-device demand.
High Bandwidth Flash targets the gap between HBM and conventional storage
Sandisk’s High Bandwidth Flash, or HBF, is designed to bring NAND closer to the performance profile required by some inference systems. The company said the architecture targets bandwidth near HBM while offering up to roughly eight times the capacity at a similar cost. Those claims will depend on the eventual configuration of servers, controllers, and software, but they illustrate the commercial appeal of using flash to relieve pressure on much scarcer HBM capacity.
HBM is optimized for moving data rapidly to and from GPUs and other accelerators, yet its capacity is limited and its cost is high. AI operators cannot simply place every dataset, retrieval index, model variant, and user session in HBM. Inference infrastructure instead needs several tiers of memory and storage, each balancing speed, capacity, power consumption, and cost.
HBF would seek a role in that middle ground. It could allow server designers to keep more data physically close to compute than traditional storage arrangements permit, reducing the need to fetch frequently used material from more distant SSD pools or network storage. That does not make NAND a replacement for HBM in bandwidth-sensitive tasks, but it could expand the range of data that remains available without paying HBM-level costs.
Sandisk is working with SK hynix on standardization for HBF, an early attempt to prevent the technology from becoming a single-vendor design. Hardware standards can determine whether a new memory category gains broad adoption, since cloud operators and server manufacturers generally need interoperable components, predictable supply, and more than one qualified supplier before redesigning systems around a new architecture.
QLC and denser NAND designs raise the capacity equation
Conventional enterprise SSDs are also changing. QLC NAND, which stores four bits in each cell, offers greater capacity per area than TLC NAND, which stores three. That difference becomes increasingly valuable in AI deployments that need massive quantities of relatively affordable storage for datasets, embeddings, logs, checkpoints, and caches.
QLC has traditionally faced trade-offs in endurance and performance, particularly in workloads involving heavy writing. Improvements in controllers, error correction, and NAND design are broadening the use cases where it can operate, especially for read-heavy or capacity-focused deployments. Inference workloads can include substantial writing through logging, caching, and changing user context, so the technology’s suitability will vary by application rather than follow a single rule.
Samsung has highlighted the density gains available through newer manufacturing approaches. The company said its BV-NAND architecture can increase storage density by about 58% compared with the prior generation. Wafer-bonding techniques, used across the memory industry, are also intended to improve capacity and performance by allowing more sophisticated layering and integration in NAND designs.
Such advances may increase the number of bits available per wafer even without major additions to fabrication capacity. That is useful for customers seeking denser drives, but it also complicates the supply outlook for producers because technology improvements can add effective supply quickly.
Long-term agreements offer suppliers clearer demand signals
Sandisk has been shifting a larger share of sales toward multi-year customer agreements rather than relying mainly on short-term transactions. According to figures reported after its latest earnings, the company had signed eight long-term agreements with six customers valued at about $93.9 billion in total, with an average duration of around four years. The agreements were said to cover roughly half of fiscal 2027 output and about two-thirds of fiscal 2028 output.
Samsung has similarly discussed a future in which long-term customer agreements account for 60% to 70% of memory sales. For NAND suppliers, committed volumes can provide more useful capacity-planning signals than spot prices, which often reflect temporary shortages or inventory adjustments rather than durable demand.
The contracts also give large cloud and enterprise buyers more confidence that storage supply will be available as they build out AI systems. That can be especially valuable when a project requires coordinated purchases of GPUs, networking equipment, servers, power infrastructure, DRAM, and SSDs. A shortage in any one component can delay an entire deployment.
The NAND cycle remains a constraint on the AI storage thesis
AI-driven inference demand may support larger and more persistent flash consumption, but NAND remains a capital-intensive and historically volatile market. Suppliers can respond to tight conditions by adding capacity, while process advances can sharply increase bit output. Those additions sometimes arrive after demand has cooled, pressuring prices and margins.
The emerging question is whether recurring inference workloads and multi-year supply agreements can make future NAND cycles less severe. They would not eliminate the risk of oversupply, especially if manufacturers expand aggressively or density gains outpace consumption. They do give flash suppliers a stronger basis for planning production around contracted demand rather than reacting mainly to volatile spot markets.
For data-center operators, the result is a more layered AI hardware strategy. GPUs and HBM remain essential for computation, but large-scale inference increasingly depends on whether systems can retain and retrieve vast quantities of data economically. NAND suppliers are betting that this requirement will make high-capacity flash a more strategic part of AI infrastructure, even if the industry’s familiar supply cycles continue.
Explore how AI complements blockchain to better understand infrastructure trends reshaping data storage, computation, and next-generation digital assets.
Disclaimer: The content on this page is provided for general informational purposes only and does not represent the views or financial advice of Toobit. We make no guarantees regarding the accuracy or completeness of this information and shall not be held liable for any errors, omissions, or outcomes resulting from its use. Investing in digital assets involves risk; users should independently evaluate their financial situation and the risks involved. For further details, please consult our Terms of Service and Risk Disclosure.
