OpenAI built a custom inference chip, benchmarked it against Nvidia's GB200 and GB300, and demonstrated it running DeepSeek R1 and Kimi K2.5 alongside its own open-weight GPT-OSS model. Then it said the chip is for internal use.

Richard Ho, who heads hardware at OpenAI, explained the apparent contradiction to Tom's Hardware without quite resolving it. Jalapeno is built for OpenAI's compute needs, he said, and then: "you could use it for anybody, honestly."

The demand argument

Ho's stated reason for keeping it inside is capacity rather than strategy.

"We have such a strong demand for compute within the company. It's going to take us a good long time to even fill our own demand, which is growing all the time," he said. "I think that we're going to have our hands full just providing compute for OpenAI for a good long time. That's not to say that it can't be used elsewhere. I believe it could be, but I think our priority is to make sure that OpenAI's compute needs are met first and foremost."

That is consistent with what the company has told investors. OpenAI's own projections put $856 billion of compute and infrastructure spending and a $278 billion cash shortfall through 2030. A company with that much appetite has an obvious use for every accelerator it can have fabricated, and no reason to sell one to somebody else.

Supply is the quieter constraint. Ho referred to "a new baseline for supply," describing two years in which he and Sam Altman toured fabs asking for capacity. He says OpenAI is in good shape internally. Having enough wafers to also serve external customers is a different problem, and it is the one that would actually decide whether Jalapeno ever ships to anyone.

Why show benchmarks at all

The interesting question is not why OpenAI is keeping the chip but why it published numbers on it. Ho says the company had not planned to show benchmarks at Hot Chips 2026 and was not certain it would present at all.

His explanation is about reputation rather than sales. "What we really wanted to demonstrate, to put to rest, the misperception in the industry that our custom inference chip was only for OpenAI models," he said. "It's programmable, and it's general purpose, and it's not hard-coded for OpenAI models."

That matters because most application-specific chips are exactly what the name says. ASICs are common and competitively performing ASICs are not, because the usual way to win on a fixed workload is to hard-wire it and lose everything else. Running DeepSeek R1 and Kimi K2.5 is the proof that Jalapeno did not take that shortcut, and the engineering anecdote supports it: a team got both models working on the chip in the two months between the A0 silicon sample and the conference presentation. Two months from first sample to running someone else's models is a claim about the toolchain more than the transistors.

The comparison set is worth noting too. The figures came from SemiAnalysis' InferenceX benchmark against the GB200 and GB300, which is Nvidia's current generation rather than a convenient older part.

Efficiency, not peak performance

Asked what drove the design, Ho was direct: "It was efficiency."

He frames efficiency as a form of compute. A modern AI data centre is constrained by power, not by how many accelerators will physically fit, so a more efficient inference engine converts the same megawatts into more useful work. On that reading, Jalapeno is not trying to beat Nvidia on throughput. It is trying to lower the cost per token inside a facility whose electrical budget is already fixed.

That is a rational thing for a buyer at OpenAI's scale to build, and it is a different product from what Nvidia sells. Nvidia's chief executive expects chip sales to double next year, which will happen whether or not OpenAI supplies part of its own inference. A custom ASIC trims the margin a supplier collects on one specific workload; it does not replace a general-purpose fleet.

The multi-generational roadmap OpenAI reiterated at Hot Chips is the part to watch. One chip is a procurement decision. A roadmap is a hardware division, and a hardware division eventually wants customers. Memory pricing is pushing in the same direction, with suppliers like Micron pressing 512GB DDR5-9200 modules at the top of the market while every large buyer looks for ways to own more of its own stack.