

OpenAI just crossed a threshold that most AI labs only dream about: it has designed and taped out its own chip. Jalapeño, OpenAI's first Intelligence Processor, is a purpose-built LLM inference accelerator co-developed with Broadcom and manufacturing partner Celestica. Engineering samples are already running ML workloads in the lab , including GPT-5.3-Codex-Spark , at production target frequency and power. This is no longer a roadmap slide. It is silicon.
What Jalapeño actually is
Inference is the phase where a trained model generates responses to real user queries , every ChatGPT answer, every Codex completion. It is fundamentally different from training, and general-purpose GPUs were not designed around it. Jalapeño is a blank-slate design for modern LLM inference, not a general-purpose accelerator adapted from earlier AI workloads.
The architecture reflects that philosophy at every level:
- The goal is to combine the power and throughput of today's leading AI accelerators with latency closer to the fastest specialized inference systems.
- The architecture reduces data movement and balances compute, memory, and networking resources to achieve realized utilization much closer to theoretical peak performance. In plain terms: most chips waste a significant fraction of their theoretical compute because data can't move fast enough. Jalapeño was designed to close that gap.
- Broadcom's silicon implementation and networking technologies, including Tomahawk networking silicon, help bring the platform to large-scale production.
- The chips are expected to feature a systolic array architecture optimized for AI inference, high-bandwidth memory (HBM, possibly HBM3E or HBM4), and manufacturing on TSMC's 3nm process node.
While OpenAI is still measuring final performance, early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art. A full technical report is expected in the coming months.
The team behind it
Jalapeño was delivered to OpenAI CEO Sam Altman and President Greg Brockman by Broadcom President and CEO Hock Tan and President Charlie Kawwas. The hardware program is led by Richard Ho, who previously led Google's TPU initiative. OpenAI hired former Lightmatter chip engineering lead and Google TPU head Richard Ho as the head of hardware.
Ho laid out the company's full-stack hardware strategy for the first time during a closed-door session at Stanford University, arguing that GPUs were never designed for the coming era of AI agents, forcing OpenAI to reclaim low-level control of AI compute through custom silicon and systems. His framing: "We're not building a chip. We're building a system."
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves
