Jalapeno, the inference accelerator OpenAI designed with Broadcom
OpenAI and Broadcom have unveiled Jalapeño, an OpenAI-designed accelerator built specifically for large language model inference and planned for gigawatt-scale deployment from the end of 2026.
OpenAI and Broadcom unveiled Jalapeño on June 24, 2026, an OpenAI-designed accelerator built specifically for large language model inference — the compute used to serve answers in ChatGPT, Codex, the API, and other AI products. The companies are targeting initial deployment by the end of 2026, but final performance numbers have not been published yet.
Definition: Jalapeño is OpenAI's first custom Intelligence Processor, designed from scratch for LLM inference rather than adapted from a general-purpose accelerator.
Example: When a user asks ChatGPT a question, Jalapeño is intended to help generate the response with less data movement and more efficient use of compute, memory, and networking resources.
Key takeaway: OpenAI is moving deeper into the infrastructure stack, but the announcement is a platform milestone — not yet a public benchmark proving superiority over existing AI accelerators.
Business impact: If the promised efficiency gains materialize, lower inference cost and latency could make high-volume AI products faster, more reliable, and cheaper to operate.
What exactly did OpenAI and Broadcom announce?
Jalapeño is the first accelerator in a multi-generation compute platform that OpenAI and Broadcom are building together. OpenAI designed the processor around its understanding of LLM fundamentals, model roadmaps, kernels, serving systems, and product requirements, while Broadcom is contributing silicon implementation, networking, connectivity, and production expertise. Celestica is supporting board, rack, and system integration.
The official OpenAI announcement describes Jalapeño as a flexible platform for current and future LLMs across the industry, not as a chip that can run only OpenAI models. Engineering samples are already running machine-learning workloads in the lab at production target frequency and power, including GPT-5.3-Codex-Spark.
Why does inference need a purpose-built chip?
Inference is the stage where a trained model generates output for a real user, so its economics are shaped by every request, token, memory transfer, and network hop. Training a model is a massive upfront workload; serving that model turns efficiency into a recurring operating cost across millions or billions of interactions.
Jalapeño is designed to optimize that serving loop directly. The companies say the architecture reduces data movement and balances compute, memory, and networking resources so realized utilization can move closer to theoretical peak performance. That matters because a chip's advertised peak is not the same as the useful throughput an AI service can achieve under real serving conditions.
OpenAI's strategy is therefore broader than replacing one processor with another. The company is co-designing the chip with the software and systems around it — including kernels, scheduling, deployment, and networking — so the full stack can be tuned for interactive LLM workloads. The same systems thinking appears in how AI agents use tools and complete multi-step tasks: product behavior depends not only on the model, but also on the infrastructure that routes and executes work.
Is Jalapeño faster than today's leading AI accelerators?
There is no public evidence yet for a definitive absolute-speed comparison. OpenAI and Broadcom say early testing indicates substantially better performance per watt than the current state of the art, but OpenAI is still measuring final performance and has promised a detailed technical report later. Until that report includes workloads, batch sizes, latency targets, power measurements, and comparison hardware, the claim should be read as an early efficiency signal rather than a completed benchmark result.
The Broadcom release also emphasizes the platform's networking and production components, including Broadcom's Tomahawk networking silicon. That focus reflects the practical reality of inference: a system can lose much of a theoretical chip advantage if memory access, interconnect bandwidth, or scheduling becomes the bottleneck.
How quickly was the chip developed?
OpenAI and Broadcom say Jalapeño moved from initial design to manufacturing tape-out in nine months. That schedule was enabled by OpenAI's engineering teams, Broadcom's implementation expertise, and the use of OpenAI models to accelerate parts of chip design and optimization.
The nine-month figure is important, but it does not mean the product is already deployed at scale. Tape-out is a manufacturing milestone; production qualification, system integration, data-center deployment, software optimization, and reliability testing still determine how quickly a chip becomes useful infrastructure. The CNBC report describes a path from prototypes in late 2026 toward larger ramp-up in 2027 and 2028, while also noting that the timeline is ambitious.
What does the partnership change for OpenAI?
Jalapeño extends OpenAI's role from model and product developer into hardware architecture and infrastructure design. That gives OpenAI more control over the trade-offs that determine how much intelligence it can serve per watt, how quickly users receive responses, and how much capacity it can deploy when demand spikes.
The move is also a capacity strategy. OpenAI says compute demand remains difficult to satisfy, and custom silicon gives the company another path alongside the general-purpose and specialized chips supplied by other infrastructure partners. A multi-generation platform could make the hardware roadmap more predictable, although it also creates execution risk: each generation must deliver enough efficiency and volume to justify the engineering investment.
What should operators watch next?
The next meaningful evidence will be the promised technical report, followed by production availability and measurements from real serving workloads. Operators should look for independent data on tokens per second, time to first token, performance per watt, memory capacity, interconnect behavior, software compatibility, and total cost of ownership rather than relying on launch language alone.
Jalapeño is already a significant strategic signal: AI companies increasingly want to shape the entire path from model architecture to deployed inference. The announcement shows where the industry is heading, but the commercial conclusion will depend on whether OpenAI and Broadcom can turn an impressive design cycle into reliable, economical infrastructure at gigawatt scale.
Frequently asked questions
What is OpenAI's Jalapeño chip?
Jalapeño is OpenAI's first custom Intelligence Processor: an accelerator designed from scratch for large language model inference, the part of an AI system that generates answers for users. OpenAI says the architecture is optimized around its models, kernels, memory movement, networking, and serving systems, while remaining flexible enough to run LLMs across the industry. It is being developed with Broadcom and Celestica as the first step in a multi-generation compute platform.
When will Jalapeño be deployed?
OpenAI and Broadcom are targeting initial deployment by the end of 2026, with the platform expanding in later generations. Broadcom CEO Hock Tan described a path from small prototype development in late 2026 toward broader ramp-up in 2027 and the first half of 2028, but those milestones are forward-looking targets rather than completed deployments.
Is Jalapeño faster than Nvidia's GPUs?
Not enough public data exists to make a fair product-to-product benchmark comparison. OpenAI and Broadcom say early testing shows substantially better performance per watt than the current state of the art, but OpenAI has not yet published final performance numbers. A detailed technical report is expected in the coming months, so claims about absolute speed, price, or competitiveness should wait for measured results.
Why is OpenAI designing its own AI chip?
OpenAI is designing more of its infrastructure to improve the cost, latency, reliability, and availability of inference at very large scale. A purpose-built accelerator can trade some general-purpose flexibility for tighter optimization around known model workloads, memory traffic, networking, and serving patterns. For OpenAI, custom silicon also reduces dependence on a single hardware path as demand for ChatGPT, Codex, API, and future agentic products continues to grow.
Alex
Founder & Lead AI Writer
Alex is the founder of Yowox and lead AI writer since 2024, breaking down complex information into clear, actionable insights for thousands of readers every day. Alex has built AI automation systems for businesses since 2024, focusing on AI agents, workflow automation, and business process optimization.
Save hours. Save thousands.
Practical guides, real workflows, and the latest AI and automation news that matters — straight to your inbox.