// TOM'S HARDWARE US — HARDWARE & GADGET
OpenAI Jalapeño design interview transcript
We sat down with OpenAI to understand how AI is being used to develop chips
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
OpenAI revealed its Jalapeño inference ASIC at Hot Chips in August 2026, a chip that leaned heavily on AI to deliver an incredibly short design window. Following the reveal, Tom's Hardware had the opportunity to sit down with the company's VP of Hardware, Richard Ho, to answer some of our most pressing questions about how the chip came into existence, future ambitions, and how AI might be used in the development of silicon.
The following article is a full transcript of our interview with Richard Ho, which has been edited for flow and clarity. You can also read additional interview transcripts we produced earlier in the year, featuring Intel, AMD, Nvidia, Valve, and more. This transcript is free to access for a limited time as part of Tom's Hardware Premium's AI Chip Design Week.
Jake Roach, Senior CPU Analyst, Tom's Hardware: It was quite the ending to Hot Chips when you dropped this. I want to start at a high level. There are a lot of reasons for OpenAI to develop its own ASIC, but was there one thing that you could point to more specifically that was a driving force? Whether it’s performance, efficiency — what was it that really drove that decision?
Richard Ho, VP Hardware, OpenAI: It is efficiency. I think that’s the main thing that we’re aiming for, because obviously, as Sam [Altman] has been saying, we are going to be compute-limited, and a compute limitation is really how much power we can get into data centers.
What we want to do is be as efficient as we can with the limited compute and limited power that we’re going to be able to get, and make the most of it. Because what we really care about is how much intelligence we can deliver to the users, and having a more efficient inference device is very useful. That’s why we focus on inference, because training happens, and you do a lot of compute with the pre-training, but really the cost to the user is on the inference side, and their perception of intelligence is going to be there. Their user experience in terms of how fast ChatGPT responds, or how fast Codex responds, or how fast the agents respond — the latency really matters.
You can see all of those things in the ingredients of what we announced. You can see that we have both a very good low-latency device for those who really care about it, and we can very easily just turn the knob and get very good throughput, so you can reduce the cost of that inference. Really, that’s the thing that we were aiming for. I’m very happy that the team managed to deliver that.
Roach: Developing your own ASIC versus going with something that’s currently on the market— were you just not satisfied with the efficiency of current offerings?
Ho: Well, I wouldn’t say that. The way to really think about it is we wanted to take advantage of the co-design opportunity that we had. It’s something that you can’t do with a third-party silicon merchant really well, because there’s a lot of research IP in the models. You just can’t share that widely because it will leak. It will get out there no matter how many NDA’s you put in place.