OpenAI and Broadcom have unveiled Jalapeño, a custom chip built specifically for LLM inference. The design targets the performance and efficiency demands of serving large AI systems at scale.

Custom inference silicon is becoming a strategic priority as model usage grows and serving costs become a constraint. Better chips can reduce latency, improve throughput, and help control infrastructure spending.

The announcement also shows OpenAI moving deeper into the hardware stack, where supply, efficiency, and optimization can shape product economics as much as model architecture.