🔥BTC/USDT

Kimi K3 increases AI hardware demand

The release of Kimi K3, a 2.8 trillion-parameter open-source large language model from Shanghai-based Moonshot AI, has intensified debate over whether more efficient AI systems will ease pressure on global computing infrastructure or make it worse. Early trading in chip stocks reflected concern that powerful Chinese open-source models could reduce demand for expensive hardware, but research from several major Wall Street firms argues that the opposite is more likely: wider access to advanced AI may increase consumption of chips, high-bandwidth memory, storage, cloud capacity, networking equipment and electricity.

Kimi K3 was unveiled on July 16 and quickly drew global attention because of its scale and benchmark performance. On the Artificial Analysis Index, the model scored 57, placing it around third to fourth globally and putting it near leading closed-source systems. It also took first place on the University of California, Berkeley’s Frontend Code Arena with 1,679 points, becoming the first open-source model to beat all overseas closed competitors on that benchmark.

The launch unsettled U.S. semiconductor shares the following day. The move echoed the market reaction seen after DeepSeek R1 appeared in early 2025, when traders worried that lower-cost Chinese AI models could weaken the case for massive infrastructure spending. This time, however, reports from UBS, Nomura, Bank of America and Citigroup have leaned toward a different conclusion. They argue that cheaper and more capable open-source models may broaden AI usage so much that total infrastructure needs keep rising.

The debate matters far beyond a single model release. Kimi K3 points to a shift in AI development toward large, multimodal systems that can process longer prompts, handle richer data and stay in reasoning mode for extended tasks. Those capabilities may reduce the cost per task, but they also raise the volume of tasks that companies, developers and consumers are likely to run.

Why Kimi K3 changed the discussion

Kimi K3 is not simply a smaller model made cheaper. Its specifications put it among the most demanding AI systems released publicly. The model has a one million-token context window, meaning it can process and retain extremely large amounts of information in a single session. It also supports constant reasoning mode and native multimodal processing across text, images and video.

Its architecture uses Kimi Delta Attention and Stable LatentMoE with 896 experts, activating 16 experts per token. Moonshot AI says this improves scaling efficiency by roughly 2.5 times compared with Kimi K2. In practical terms, the design aims to deliver stronger performance without making each task proportionally more expensive.

That combination is central to the market debate. A model that is both powerful and more accessible may encourage more companies to integrate advanced AI into software, customer service, coding, design, analytics and media workflows. Even if each AI request becomes cheaper, total token usage can rise sharply as more people use the tools more often.

Citigroup’s Peter Lee referred to this as a version of the Jevons paradox, the economic idea that improvements in efficiency can increase overall resource consumption rather than reduce it. When a technology becomes cheaper and easier to use, demand often expands enough to offset the efficiency gains.

Pricing shows a mid-market strategy

Kimi K3’s pricing also suggests that Moonshot AI is not positioning the model as a bare-bones budget product. Input costs are listed at $3 per million tokens, cached inputs at $0.30 per million tokens and output at $15 per million tokens. The average task cost has been estimated at about $0.94.

That is below some higher-priced proprietary models cited in market reports, including Claude Fable 5 at $2.75 and Claude Opus 4.8 at $1.80. It is close to GPT-5.6 Sol at about $1.04, while remaining above lower-priced models such as GLM-5.2 at $0.32 to $0.47 and DeepSeek V4 Pro at $0.04.

The pricing structure indicates a focus on high capability at moderate cost rather than an aggressive discount strategy. For enterprises and developers, that could make Kimi K3 attractive for tasks that require long context, multimodal inputs or sustained reasoning but still need to stay within budget.

For hardware suppliers, the details matter. Longer prompts and larger output workloads can increase demand for memory, storage and fast interconnects. Cached inputs may reduce repeated computation, but they also require systems to store and retrieve more data efficiently.

Hardware demand may shift, not shrink

The immediate market concern after Kimi K3’s debut was that efficient open-source models could slow purchases of advanced AI hardware. Wall Street research has pushed back against that view.

Nomura’s Asia-Pacific technology team said competition around generative AI and progress toward more advanced reasoning systems are likely to support continued heavy spending on infrastructure. Bank of America’s Vivek Arya said U.S. AI companies may respond to stronger Chinese competition by developing even larger and faster models, which would require more computing capacity rather than less.

UBS analysts highlighted Kimi K3’s open-source nature as especially important. Open-source models can spread quickly through developer communities, cloud platforms and enterprise software stacks. Once a model is widely available, more teams can modify, host and deploy it. That creates demand across a broader base of users instead of concentrating usage inside a few proprietary AI platforms.

The hardware burden also changes with the type of workload. Training giant models is one major source of demand, but inference — the process of running models for end users — becomes increasingly important as adoption grows. Kimi K3’s long context window and multimodal inputs could make inference more memory-heavy and storage-heavy than shorter, text-only workloads.

Citigroup’s Lee noted that longer context and larger cache usage can increase demand for DDR5 memory and enterprise solid-state drives. UBS also pointed to high-bandwidth memory and distributed storage systems as important for Kimi K3-class workloads.

Memory and storage move into focus

Among hardware categories, memory and storage may see some of the clearest benefits if the market shifts toward long-context AI systems. AI models that process larger documents, video files, code bases and enterprise databases need fast access to more data. That requires not only graphics processors, but also memory close to the processor and storage systems capable of feeding data at high speed.

UBS estimated that free cash flow for memory producers could reach 30% of market capitalization by 2028, with Micron projected at 47%. Citigroup and Nomura maintained positive views on major memory makers, citing tight global supply and rising demand from AI workloads.

The storage effect could be especially important. As models use longer context windows, companies may cache more prompts, documents, embeddings, user histories and task outputs. That raises demand for enterprise solid-state drives, data management systems and distributed storage architectures.

Large-scale AI clusters also need the ability to move data quickly between chips and servers. Network equipment suppliers may benefit from the need for denser interconnects, faster switching and lower-latency communication across AI clusters. As models become more complex, the weakest link is often not the processor alone, but the full system that connects compute, memory, storage and networking.

Chips remain central to the buildout

Graphics processors and AI accelerators remain at the center of the infrastructure story. Citigroup reported that deployment of a single Kimi K3 supernode requires more than 64 GPUs. That scale suggests that even efficient open-source models can be difficult and expensive to host at high performance.

Nomura reaffirmed positive ratings for leading fabrication and advanced packaging companies, pointing to the role they play in delivering higher-performing AI chips. The report also cited efficiency gains in NVIDIA’s GB300 NVL72 systems, which are described as up to 25 times better than the prior Hopper generation for certain workloads. Such improvements could support Kimi K3-style models by making each cluster more productive, but may also encourage greater use.

This is the key tension for traders watching the semiconductor sector. More efficient chips and more efficient models can reduce cost per unit of output. But if demand for AI output rises faster than efficiency improves, total hardware spending can still climb.

Bank of America added a note of caution, saying infrastructure growth could cool temporarily if efficiency gains significantly outpace growth in model usage. For now, available data on token processing and enterprise AI adoption point toward expansion rather than contraction.

Cloud platforms and data centers gain importance

Cloud infrastructure is another likely beneficiary of the open-model ecosystem. Many companies do not want to buy, configure and manage their own AI clusters. Instead, they rely on cloud providers and hosting platforms that can offer multiple models, flexible pricing and scalable capacity.

Open-source models such as Kimi K3 strengthen the case for multi-model hosting services. Developers may want to compare proprietary and open systems, route tasks to different models and manage costs across workloads. That makes cloud platforms more important as intermediaries between model creators and business users.

Nomura identified several Asian data center operators expanding capacity to support the open-model ecosystem. This expansion reflects a broader regional trend: China-based and Asia-based AI companies are becoming more competitive in model performance, while demand for local hosting, compliance and lower-latency services continues to grow.

OpenRouter data shows China’s share of global AI token usage rising from less than 2% a year ago to more than 45%. That shift indicates that Chinese models and Chinese AI platforms are gaining significant usage share, though token data can vary depending on platform coverage and methodology.

Bank of America estimates that 55% of U.S. companies now subscribe to AI tools or platforms, with corporate users spending an average of $4,833 per employee per month on AI services. If spending continues to rise, infrastructure pressure is likely to remain significant even as individual model calls become cheaper.

Energy demand becomes a harder constraint

The push to build larger AI clusters is not only a hardware issue. It is also becoming an energy issue. High-density data centers require stable power, cooling systems, backup generation and upgraded grid connections. In many regions, power availability is becoming one of the biggest obstacles to new data center development.

Deloitte has projected that data center power demand in the United States could rise from about four gigawatts to 123 gigawatts by 2035. The International Energy Agency has also noted growing interest in power agreements tied to small modular nuclear reactors, with such deals rising from 25 gigawatts to 45 gigawatts.

These numbers show how quickly digital infrastructure is becoming linked to energy policy. Data centers cannot expand indefinitely without transmission upgrades, local grid planning and new power sources. In areas where electricity systems are already stretched, rapid data center growth raises the risk of bottlenecks, delays and political resistance.

Cooling is another practical constraint. Advanced AI servers generate enormous heat and often require specialized liquid cooling or other thermal management systems. That adds cost and complexity for data center operators and increases the importance of site selection.

Digital infrastructure tokens draw renewed attention

The strain on centralized computing and energy infrastructure has also revived interest among crypto traders in decentralized physical infrastructure networks, often called DePIN. These projects aim to coordinate real-world resources such as computing power, wireless coverage, storage, mapping data or sensor networks through blockchain-based systems.

However, the sector remains volatile and highly speculative. Tokens linked to computing power or distributed storage have moved sharply with broader crypto conditions, changes in AI sentiment and project-specific adoption metrics. Because many of these networks are still early, token prices can move far faster than real usage.

The most important question for traders is whether networks tied to physical infrastructure can show sustained demand from paying users. In the case of distributed GPU projects, that means proving that developers and companies can reliably rent useful compute at competitive prices. For storage networks, it means showing that customers are storing meaningful data and renewing usage over time.

AI growth may create a stronger narrative for these projects, but narrative alone is not enough. Real server capacity, uptime, network performance, customer demand and transparent revenue will matter more as the sector matures. Projects without strong real-world usage may struggle even if AI infrastructure shortages intensify.

A split between open and closed ecosystems

The global AI market is increasingly developing along two tracks. Open-source models such as Kimi K3 are gaining strength in midrange, enterprise and cost-conscious markets where customization and deployment flexibility matter. U.S. proprietary labs continue to focus on frontier systems, advanced computation and research-heavy workloads that require very large budgets and specialized infrastructure.

These two tracks are not necessarily in conflict from a hardware-demand perspective. Both can increase resource use. Open models can spread AI usage to more developers and businesses, while frontier labs can push the high end of compute demand higher. If both trends continue, the result may be more pressure across the entire infrastructure stack.

Short-term stock volatility is likely to continue as traders assess whether each new model release helps or hurts specific companies. But the broader pattern points toward rising usage, longer context windows, larger data flows and deeper AI integration into business software.

Kimi K3’s main market impact may therefore be less about one day of semiconductor trading and more about what it signals for the next stage of AI adoption. Efficient, powerful and accessible models can lower barriers to entry, but they can also increase the total amount of computation the world consumes. For chipmakers, memory suppliers, storage companies, cloud platforms, data center operators and power providers, that may keep demand under pressure for years.


As AI reshapes infrastructure demand, explore crypto’s evolution in AI in banking and its impact on digital markets.

Disclaimer: The content on this page is provided for general informational purposes only and does not represent the views or financial advice of Toobit. We make no guarantees regarding the accuracy or completeness of this information and shall not be held liable for any errors, omissions, or outcomes resulting from its use. Investing in digital assets involves risk; users should independently evaluate their financial situation and the risks involved. For further details, please consult our Terms of Service and Risk Disclosure.

Sign up and trade to earn over 15,000 USDT
Sign up