OpenAI’s first in-house AI inference chip, code-named Jalapeño, outperformed Nvidia’s GB200 and GB300 systems on throughput, latency and energy efficiency in third-party tests conducted by semiconductor research firm SemiAnalysis, according to results from an engineering-sample system tested at OpenAI’s lab.
The results put OpenAI closer to becoming a large-scale designer and operator of its own AI hardware, a role that could reduce its dependence on merchant GPUs for serving models. Jalapeño was developed with Broadcom for inference, the process of generating responses from trained AI models, and SemiAnalysis said production use could begin later this year. Volume manufacturing is expected to ramp in 2027.
SemiAnalysis tested Jalapeño with its InferenceX benchmark suite and said the chip led every processor it had previously measured from Nvidia, AMD and Google on performance per watt. The firm said competing systems were configured using their best-known settings, while Jalapeño achieved its results without speculative decoding or splitting prompt processing and token generation across separate hardware pools.
Jalapeño leads on GPT-OSS throughput and latency
On the GPT-OSS 120B model, SemiAnalysis measured Jalapeño at roughly 1,459 tokens per second, compared with 535 tokens per second for Nvidia’s GB200. A token is a unit of text processed by an AI model; higher token throughput generally allows an operator to serve more requests from the same hardware.
The reported gap was especially large in one interactive “8k1k” workload, which uses an 8,000-token input and a 1,000-token output. SemiAnalysis said Nvidia’s GB300 reached a maximum decoding speed of 169 tokens per second in that scenario, while Jalapeño delivered 104.3 times more throughput at that interactive operating point.
Latency results also favored the OpenAI design. SemiAnalysis measured end-to-end latency of 1.65 seconds for Jalapeño, compared with nearly six seconds for Nvidia’s GB300, a difference of about 3.6 times. Lower latency is particularly valuable for consumer chat products, coding tools and AI agents, where delays can compound across multi-step tasks.
The testing does not establish Jalapeño as a universal replacement for Nvidia hardware. SemiAnalysis said the chosen GPT-OSS model was not among the most advanced models available, and long-context, multi-turn AgentX testing had not yet been completed on Jalapeño. Nvidia and AMD have published AgentX results on larger models, including DeepSeek V4 Pro and Kimi K3, according to the research firm.
Energy results target the largest AI operating expense
SemiAnalysis reported that Jalapeño generated about 53 million GPT-OSS tokens per megawatt per second, against about 10 million tokens per megawatt per second for Nvidia’s GB200 NVL72. The metric is effectively tokens produced per joule of energy and places power consumption near the center of OpenAI’s chip strategy.
Energy efficiency has become a practical constraint on AI expansion as model providers add more racks, secure grid capacity and build cooling infrastructure. A chip that produces more usable inference output per watt would allow an operator to fit more AI serving capacity within a fixed power allocation, although the final economics will depend on model size, utilization rates, software maturity and data-center design.
SemiAnalysis estimated a system-level total cost of ownership of about $1.56 per Jalapeño chip per hour, including power delivery, cooling and networking. It placed Nvidia’s H100 at $1.55 per chip per hour and Nvidia’s forthcoming Vera Rubin platform at $3.61 per chip per hour.
The comparison comes with an important generational caveat. SemiAnalysis said Vera Rubin is the closer architectural peer because both platforms use HBM4 high-bandwidth memory. It estimated Rubin’s performance per watt at about 5.4 times that of the GB200 NVL72 and described Jalapeño and Rubin as broadly close on per-token total cost of ownership.
OpenAI used AI tools to develop the chip
SemiAnalysis said Jalapeño moved from design start to tape-out in roughly 16 months, faster than the 18-to-36-month range it described as common across the industry. The crucial CoWoS advanced-packaging tape-out was completed in November 2025, about nine months before the reported test results.
OpenAI used internal AI systems during development, including GPT-Astra and an internal version of Codex for kernel engineering, according to SemiAnalysis. The firm said AI-assisted design reduced SIMD-unit area by 8% and matrix-engine area by 10%, while also improving timing and power measurements relative to the initial chip design.
For software, SemiAnalysis said Codex wrote Jalapeño kernels of around 3,000 lines. In tests covering attention and mixture-of-experts modules, it reported that AI-generated code ran 1.5 to 1.8 times faster than versions produced by top human engineers.
That claim points to a potential advantage beyond the silicon itself. AI chips often gain substantial performance through kernel tuning, compilers and inference engines after the hardware reaches the lab. SemiAnalysis reported that Jalapeño’s throughput at one interactive speed more than doubled in less than two weeks, while support for tensor parallelism expanded from TP8 to TP32 in eight days. That expansion would allow the system to spread models across larger groups of chips, including cross-rack deployments.
A design aimed at mixed AI traffic
Jalapeño is designed as a general-purpose inference processor rather than a narrow accelerator for one model configuration, SemiAnalysis said. Its architecture connects compute cores directly to HBM memory slices and uses a dedicated high-bandwidth collective network to synchronize cores.
The chip uses out-of-order execution cores and L1 cache rather than the software-managed scratchpads found in many AI accelerators. SemiAnalysis said Jalapeño provides 15.4 TB/s of HBM4 bandwidth per package and estimated HBM bandwidth per watt at 22, compared with 11.1 for Rubin and 5.71 for GB300.
OpenAI has also avoided prefill/decode disaggregation, an approach that assigns prompt processing and text generation to different resource pools. SemiAnalysis said Jalapeño instead maintains a homogeneous pool of hardware that can shift capacity between latency-sensitive queries and high-throughput batch work as traffic changes.
The publicly tested chip was an A0 stepping, or early silicon revision. SemiAnalysis said a B0 version has entered wafer fabrication and is expected to improve performance per watt by about 25%, while retaining a 700-watt thermal design power. It reported B0 compute-die performance of 13.4 PFLOPs for MXFP4 operations.
Rack-scale ambitions accompany the chip launch
OpenAI and Celestica designed a rack containing 128 Jalapeño chips, according to SemiAnalysis. A two-rack system draws about 160 kilowatts, which the firm compared with the power level of an Nvidia GB300 dual-width rack.
A single scale-out networking domain would connect up to 2,048 Jalapeño processors across 16 racks. SemiAnalysis also cited a 100-megawatt deployment milestone and reported that OpenAI and Broadcom have signed a 10-gigawatt custom-accelerator agreement.
For cryptocurrency markets, the immediate relevance is limited to projects that genuinely purchase, deploy or rent AI compute. Benchmark gains for a proprietary OpenAI chip do not by themselves create a basis for valuing tokens tied to decentralized computing, storage or AI-agent narratives. The more concrete near-term effect would be increased pressure on GPU suppliers, data-center operators and cloud platforms to show how their hardware costs and power use compare as OpenAI begins bringing custom inference capacity into production.
Exploring AI hardware breakthroughs like Jalapeño? Learn how AI copy trading applies advanced algorithms to real trading decisions.
Disclaimer: The content on this page is provided for general informational purposes only and does not represent the views or financial advice of Toobit. We make no guarantees regarding the accuracy or completeness of this information and shall not be held liable for any errors, omissions, or outcomes resulting from its use. Investing in digital assets involves risk; users should independently evaluate their financial situation and the risks involved. For further details, please consult our Terms of Service and Risk Disclosure.
