OpenAI has published the first benchmark results for Jalapeño, its custom chip for running rather than training AI models. Across three public models, the company reports 1.5 to 1.9 times more inference work per watt and 1.7 to 3.6 times lower end-to-end latency than the commercial systems used for comparison.
The tests covered GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T on SemiAnalysis’ public InferenceX benchmark. For highly interactive workloads, OpenAI reports a 2.1 to 4.1 times performance advantage. Jalapeño is rated at 700 watts, though measured sustained draw stayed at or below 550 watts in the tested workloads.
The chip, memory, network and serving software were designed together to reduce movement of model state. OpenAI says the architecture keeps the key-value cache—the temporary memory used while generating a response—local and adjusts compute, memory and networking resources for different inference phases.
These are OpenAI’s own early measurements, not broad independent deployment results. The company says Jalapeño is working first-party silicon and plans to ramp it over the coming months, but it has not provided customer availability or pricing details.