OpenAI unveils Jalapeño: its inference chip speeds up model responses by nearly four times

19 September 202617 views

Developed together with Broadcom, the Jalapeño processor promises a notable gain in performance per watt and lower latency when running large language models — in interactive scenarios, a speedup of up to 4.1x is claimed. Deployment in OpenAI's infrastructure is planned for the end of the year, though the company refuses to write off Nvidia's accelerators.

OpenAI unveils Jalapeño: its inference chip speeds up model responses by nearly four times

What OpenAI unveiled on August 25

OpenAI presented the test results of its own AI processor — Jalapeño. This is not a third-party custom-built design, but the first chip brought to the stage of measurable benchmarks: the project was carried out together with Broadcom, and the public report appeared on August 25, 2026.

Sam Altman commented on the news on X with a single phrase — "we made a chip and it is fast." Brief, but to the point: the main emphasis is not on flexibility or versatility, but on raw response speed.

The key detail is the intended purpose. The chip is designed for inference of large language models, that is, for serving already-trained models at the moment a user asks a question. General AI workloads, training, and architecture experiments remain on other hardware. OpenAI plans to deploy Jalapeño in its own compute infrastructure by the end of 2026.

Why the focus is specifically on inference

The difference between training a model and running it in production is the difference between constructing a building and operating it. Once the model is built, costs and delays shift elsewhere: how quickly the model's data reaches the compute blocks, how many times it has to be moved around, and how much time is spent on the exchange between memory and network components.

The Jalapeño developers were solving exactly this problem — the misalignment of compute, memory, and network. The idea is to keep the model's parameters and KV cache data as close as possible to where active processing happens. The fewer the movements of information between system layers, the shorter the path from request to first token and the smoother the response on long generations. The company describes this as tighter coordination of all three layers — compute, memory, and network — instead of optimizing them independently.

The second claimed effect is more useful AI work per watt of energy at peak load. For data centers, this is not an abstract metric: electricity and cooling have long been a constraint on scaling.

How and on what it was tested

Models and benchmark

The tests ran on three models of different scales: GPT OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Measurements were taken via InferenceX — a public benchmark maintained by SemiAnalysis. The comparison was against commercial AI systems, that is, not against an abstract reference point, but against working solutions on the market.

The numbers

  • at peak throughput — 1.5 to 1.9 times more AI work per watt;
  • end-to-end latency dropped by 1.7–3.6 times;
  • in interactive scenarios, a performance increase of 2.1–4.1 times was recorded — this is exactly where the phrasing "almost four times faster" came from.

Where the gains are most noticeable

The spread in the numbers is not measurement error but an important signal. The difference between the lower and upper bounds depends on the nature of the workload. Where the system runs in batches and is loaded to the max, the gain is more modest. Where requests come in one at a time and the user is waiting for a response in real time, the advantage becomes maximal.

This is also why there is particular interest in agents. When an AI system performs a multi-step task — planning, calling tools, retrieving data, generating a response — the delays add up. Each extra step adds waiting time, and over a long chain the difference between "one and a half seconds" and "half a second" turns into a completely different user experience. Reducing latency in interactive scenarios directly addresses this problem.

OpenAI expects such figures to support faster products and improve the economics of infrastructure: it is cheaper to serve the same volume of requests.

Its own chips — but not instead of Nvidia

Here it is important not to mistake the signal. The company is not planning to discard Nvidia accelerators: they remain in the loop for both training and inference. Jalapeño is positioned as an additional compute layer for a specific class of tasks, not as an immediate replacement for existing hardware.

An interesting detail from the company's own account — the development cycle was shortened to nine months, and OpenAI's own models played a notable role in this. In other words, AI helped design the chip on which AI will then run.

What's next

Gen 2 is already in development. This indicates that this is not a one-off experiment, but a long-term bet on its own silicon as part of the future AI infrastructure.

One caveat is worth keeping in mind: all the figures given are claimed results on a specific public benchmark, not an independent verification by outside hands. Much will become clearer once the chip is actually deployed in infrastructure by the end of 2026 and data from real production workloads, rather than test runs, becomes available. For now, the main takeaway is simple: the race for model response speed has moved from the software level to the hardware level.

Frequently asked questions

Related materials

All materials
OpenAI unveils Jalapeño: its inference chip speeds up model responses by nearly four times