toobit
Buy crypto
Buy cryptoThe fastest path to your first trade
P2P tradingTrade at the best prices with multiple local payment options
Bank cardPay with Visa or Mastercard
Third-partyPay via MoonPay, Advcash, Simplex, and more
DepositTransfer from another wallet
Markets
OpportunitiesTrack market sentiment and top movers
OverviewReal-time prices for all trading pairs
Futures
USDT-M PerpetualContracts settled in USDT
USDC-M PerpetualContracts settled in USDC
Event ContractsTrade on the outcome of market events
Prediction MarketTurn insights into value
Lite PerpetualSimple contracts made for easy trading
Demo TradingPractice trading in a risk-free environment
Trading BotsAutomated grid and DCA strategies
TradFi
Trade
SpotBuy and sell cryptocurrencies
DEX +Trade popular on-chain Web3 tokens in seconds
LaunchpadAccess early-stage token listings
ConvertZero-fee instant asset swaps
API TradingAutomate trading strategies with custom scripts and apps
Toobit SynapseMarket insights driven by AI analysis
Toobit x TradingViewTrade directly from TradingView charts
Agent Trade KitEquip AI agents with trading and account skills
Rewards
Copy
Follow Lead TradersCopy trades from top-performing profiles
Be a Lead TraderShare your trades and earn commissions
More
Finance
EarnPut your idle assets to work
Partnerships
Broker ProgramMonetize API volume and trading infrastructure
Ambassador ProgramRepresent the exchange and earn monthly incentives
Toobit x Nova.MemeLaunch and trade memecoins with instant liquidity
Learn
AcademyTechnical analysis and crypto trading guides
Support CenterSelf-service help and 24/7 technical assistance
Announcement CenterLatest listings, campaigns, and official product news
NewsBreaking crypto news and market moves
BlogMarket insights and exchange updates
Explore
Toobit VIP ProgramEnjoy fee discounts and many exclusive rewards.
InsightsStay updated on the latest crypto news
Toobit CommunityConnect with The Hive, our global community of traders
3 years togetherCelebrate our journey and the community that built it
About usThe story behind the award-winning exchange
Suggestions & FeedbackShare your ideas to improve the exchange
Proof of ReservesTrust built on 100% reserves
Log in
Sign up
🔥BTC/USDT
Scan to download
iOS or Android version app
More download options

OpenAI Jalapeño inference chip beats Nvidia benchmarks

2026-08-26 10:15

OpenAI’s first in-house AI inference chip, code-named Jalapeño, outperformed Nvidia’s GB200 and GB300 systems on throughput, latency and energy efficiency in third-party tests conducted by semiconductor research firm SemiAnalysis, according to results from an engineering-sample system tested at OpenAI’s lab.

The results put OpenAI closer to becoming a large-scale designer and operator of its own AI hardware, a role that could reduce its dependence on merchant GPUs for serving models. Jalapeño was developed with Broadcom for inference, the process of generating responses from trained AI models, and SemiAnalysis said production use could begin later this year. Volume manufacturing is expected to ramp in 2027.

SemiAnalysis tested Jalapeño with its InferenceX benchmark suite and said the chip led every processor it had previously measured from Nvidia, AMD and Google on performance per watt. The firm said competing systems were configured using their best-known settings, while Jalapeño achieved its results without speculative decoding or splitting prompt processing and token generation across separate hardware pools.

Jalapeño leads on GPT-OSS throughput and latency

On the GPT-OSS 120B model, SemiAnalysis measured Jalapeño at roughly 1,459 tokens per second, compared with 535 tokens per second for Nvidia’s GB200. A token is a unit of text processed by an AI model; higher token throughput generally allows an operator to serve more requests from the same hardware.

The reported gap was especially large in one interactive “8k1k” workload, which uses an 8,000-token input and a 1,000-token output. SemiAnalysis said Nvidia’s GB300 reached a maximum decoding speed of 169 tokens per second in that scenario, while Jalapeño delivered 104.3 times more throughput at that interactive operating point.

Latency results also favored the OpenAI design. SemiAnalysis measured end-to-end latency of 1.65 seconds for Jalapeño, compared with nearly six seconds for Nvidia’s GB300, a difference of about 3.6 times. Lower latency is particularly valuable for consumer chat products, coding tools and AI agents, where delays can compound across multi-step tasks.

The testing does not establish Jalapeño as a universal replacement for Nvidia hardware. SemiAnalysis said the chosen GPT-OSS model was not among the most advanced models available, and long-context, multi-turn AgentX testing had not yet been completed on Jalapeño. Nvidia and AMD have published AgentX results on larger models, including DeepSeek V4 Pro and Kimi K3, according to the research firm.

Energy results target the largest AI operating expense

SemiAnalysis reported that Jalapeño generated about 53 million GPT-OSS tokens per megawatt per second, against about 10 million tokens per megawatt per second for Nvidia’s GB200 NVL72. The metric is effectively tokens produced per joule of energy and places power consumption near the center of OpenAI’s chip strategy.

Energy efficiency has become a practical constraint on AI expansion as model providers add more racks, secure grid capacity and build cooling infrastructure. A chip that produces more usable inference output per watt would allow an operator to fit more AI serving capacity within a fixed power allocation, although the final economics will depend on model size, utilization rates, software maturity and data-center design.

SemiAnalysis estimated a system-level total cost of ownership of about $1.56 per Jalapeño chip per hour, including power delivery, cooling and networking. It placed Nvidia’s H100 at $1.55 per chip per hour and Nvidia’s forthcoming Vera Rubin platform at $3.61 per chip per hour.

The comparison comes with an important generational caveat. SemiAnalysis said Vera Rubin is the closer architectural peer because both platforms use HBM4 high-bandwidth memory. It estimated Rubin’s performance per watt at about 5.4 times that of the GB200 NVL72 and described Jalapeño and Rubin as broadly close on per-token total cost of ownership.

OpenAI used AI tools to develop the chip

SemiAnalysis said Jalapeño moved from design start to tape-out in roughly 16 months, faster than the 18-to-36-month range it described as common across the industry. The crucial CoWoS advanced-packaging tape-out was completed in November 2025, about nine months before the reported test results.

OpenAI used internal AI systems during development, including GPT-Astra and an internal version of Codex for kernel engineering, according to SemiAnalysis. The firm said AI-assisted design reduced SIMD-unit area by 8% and matrix-engine area by 10%, while also improving timing and power measurements relative to the initial chip design.

For software, SemiAnalysis said Codex wrote Jalapeño kernels of around 3,000 lines. In tests covering attention and mixture-of-experts modules, it reported that AI-generated code ran 1.5 to 1.8 times faster than versions produced by top human engineers.

That claim points to a potential advantage beyond the silicon itself. AI chips often gain substantial performance through kernel tuning, compilers and inference engines after the hardware reaches the lab. SemiAnalysis reported that Jalapeño’s throughput at one interactive speed more than doubled in less than two weeks, while support for tensor parallelism expanded from TP8 to TP32 in eight days. That expansion would allow the system to spread models across larger groups of chips, including cross-rack deployments.

A design aimed at mixed AI traffic

Jalapeño is designed as a general-purpose inference processor rather than a narrow accelerator for one model configuration, SemiAnalysis said. Its architecture connects compute cores directly to HBM memory slices and uses a dedicated high-bandwidth collective network to synchronize cores.

The chip uses out-of-order execution cores and L1 cache rather than the software-managed scratchpads found in many AI accelerators. SemiAnalysis said Jalapeño provides 15.4 TB/s of HBM4 bandwidth per package and estimated HBM bandwidth per watt at 22, compared with 11.1 for Rubin and 5.71 for GB300.

OpenAI has also avoided prefill/decode disaggregation, an approach that assigns prompt processing and text generation to different resource pools. SemiAnalysis said Jalapeño instead maintains a homogeneous pool of hardware that can shift capacity between latency-sensitive queries and high-throughput batch work as traffic changes.

The publicly tested chip was an A0 stepping, or early silicon revision. SemiAnalysis said a B0 version has entered wafer fabrication and is expected to improve performance per watt by about 25%, while retaining a 700-watt thermal design power. It reported B0 compute-die performance of 13.4 PFLOPs for MXFP4 operations.

Rack-scale ambitions accompany the chip launch

OpenAI and Celestica designed a rack containing 128 Jalapeño chips, according to SemiAnalysis. A two-rack system draws about 160 kilowatts, which the firm compared with the power level of an Nvidia GB300 dual-width rack.

A single scale-out networking domain would connect up to 2,048 Jalapeño processors across 16 racks. SemiAnalysis also cited a 100-megawatt deployment milestone and reported that OpenAI and Broadcom have signed a 10-gigawatt custom-accelerator agreement.

For cryptocurrency markets, the immediate relevance is limited to projects that genuinely purchase, deploy or rent AI compute. Benchmark gains for a proprietary OpenAI chip do not by themselves create a basis for valuing tokens tied to decentralized computing, storage or AI-agent narratives. The more concrete near-term effect would be increased pressure on GPU suppliers, data-center operators and cloud platforms to show how their hardware costs and power use compare as OpenAI begins bringing custom inference capacity into production.


Exploring AI hardware breakthroughs like Jalapeño? Learn how AI copy trading applies advanced algorithms to real trading decisions.

Disclaimer: The content on this page is provided for general informational purposes only and does not represent the views or financial advice of Toobit. We make no guarantees regarding the accuracy or completeness of this information and shall not be held liable for any errors, omissions, or outcomes resulting from its use. Investing in digital assets involves risk; users should independently evaluate their financial situation and the risks involved. For further details, please consult our Terms of Service and Risk Disclosure.

About
About us
Terms of Use
Privacy Policy
Risk disclosure
Toobit Community
Announcement Center
Security solutions
Toobit Shield
Proof of Reserves
Services
Trade
Futures
Copy
Affiliate Program
API
Listing application
Bug bounty
Support
Support Center
Academy
Referral
Fee rate policy
Official verification
Network monitoring
Suggestions & Feedback
Buy crypto
Buy Bitcoin
Buy Ethereum
Buy Dogecoin
Buy TON
Buy SOL
Buy XRP
Contact
Customer Support
support@toobit.com
Business
listing@toobit.com
Overview
market@toobit.com
Legal
legal@toobit.com
Apps
Google Play
App Store
Android APK
Community
TwitterMediumYoutubeDiscordRedditFacebookCoinMarketCapCoinCodexCoinGeckoLinkedinQuoraThreads
Download app
Warning

© 2026 Toobit.com. All rights reserved.