On August 25, OpenAI published the first test results for its custom Jalapeño inference chip. This is a material story within the latest 72-hour extension, not an event first announced on August 27. The performance figures are primarily OpenAI's measurements, not a completed independent chip review.
What numbers did OpenAI publish?
OpenAI says that on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T tests, Jalapeño delivered about 1.5x to 1.9x the AI work per watt of NVIDIA GB300 and 1.7x to 3.6x lower end-to-end latency. For interactive workloads, OpenAI reports performance gains of about 2.1x to 4.1x.
The tests used SemiAnalysis InferenceX and normalized comparison-chip power using published ratings. OpenAI also says Jalapeño is rated at 700W but measured sustained power stayed at or below 550W in the tested workloads. Those conditions directly affect the comparison.
How should the results be read?
First, this is a first-party benchmark published by the company, not a comprehensive independent validation. It shows what OpenAI is optimizing—end-to-end inference latency, work per watt and serving cost—but batch size, input and output length, software stack and power definitions can change the outcome.
Second, Jalapeño is not a consumer product that people can buy and install. OpenAI did not announce a general developer supply path or price in this release. Any claim of beating NVIDIA applies to OpenAI's stated test conditions, not every AI workload.
Why is OpenAI building an inference chip?
Large AI services pay not only for training but also for continuous inference requests. A chip that does more work at the same power or lowers latency at the same workload could improve data-center capacity and operating costs. That is why Jalapeño emphasizes work per watt and end-to-end latency.
OpenAI says Jalapeño took about nine months from AI-assisted design to tape-out. That shows models can participate in hardware design, but it does not mean the chip is in mass production or that all design and verification work can be automated.
What is confirmed so far
Jalapeño matters because it moves OpenAI from simply consuming GPUs toward designing inference infrastructure. For now, the precise description is that OpenAI published a set of self-measured results. Broader reproducibility will require more transparent testing and outside validation.
