OpenAI plans to use its Jalapeño inference chip to meet its own growing compute demand first, hardware vice president Richard Ho said in an interview published September 28, 2026. Ho said wider use could be possible, but the company’s needs take priority.
OpenAI puts its own compute needs first
Jalapeño is a custom application-specific integrated circuit, or ASIC, designed for AI inference: the computing work that runs a trained model to produce responses. Ho described efficiency as the chip’s main design goal.
He said OpenAI’s demand for compute is growing and that meeting it would take priority. Ho left open the possibility of using Jalapeño elsewhere, but framed that as a possibility rather than a customer rollout.
A programmable inference chip demonstrated on several models
Ho described Jalapeño as programmable and not hard-coded for OpenAI models. Public demonstrations included GPT-OSS, DeepSeek R1 and Kimi K2.5. Those demonstrations show the chip running these models; they do not establish identical performance across models.
The public InferenceX test setup
The public InferenceX configuration discussed by Ho used 8,000 input tokens and 1,000 output tokens. That gives the benchmark a specific workload: a fixed-length prompt and response, rather than a measure of every possible inference task.
The design effort and production plan Ho described
Ho said Jalapeño’s design began from scratch in November 2025 and took nine months to reach tape-out, the point when a chip design is sent for manufacturing. He said AI tools assisted the engineering work.
In the September 2026 interview, Ho described a plan to place a very small number of chips in a production environment for testing, if possible, before a more substantial production ramp in 2027.