OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom AI inference processor, on June 24, 2026. It is the first named chip from a strategic collaboration announced in October 2025, which aims to deploy 10 gigawatts of OpenAI-designed accelerators and networking systems. Initial deployment is planned for the end of 2026, but the companies have not disclosed public benchmarks, pricing, manufacturing details, or a way for outside customers to buy or rent the chip.
What OpenAI and Broadcom announced
The announcement unfolded in two stages. On October 13, 2025, OpenAI and Broadcom announced a multiyear collaboration to develop and deploy 10 gigawatts of custom AI accelerators and networking systems. On June 24, 2026, they unveiled the first processor from the effort, named Jalapeño. The chip unveiling was not the original signing of the partnership.
The original plan called for deployment to begin in the second half of 2026 and continue through the end of 2029. The companies now say initial Jalapeño deployment is planned by the end of 2026. Those dates are targets, not evidence that systems have already been installed at scale. OpenAI’s October 2025 announcement describes the broader infrastructure plan; its June 2026 announcement names the first processor.
What Jalapeño is designed to do
OpenAI and Broadcom describe Jalapeño as an LLM-optimized inference processor and OpenAI’s first “Intelligence Processor.” Inference is the work of running a trained model to respond to prompts, generate code or media, or carry out tasks. Training, by contrast, is the process of building or refining a model.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
The stated focus is serving models efficiently at scale, rather than replacing every kind of accelerator OpenAI uses. A processor specialized for inference could, in principle, suit recurring model-serving workloads particularly well. Whether that specialization delivers lower costs or better service in practice depends on the chip, software, memory, networking, utilization, and the systems around it.
Who is responsible for each part?
- OpenAI designed the accelerator around its models, kernels, serving systems, and product requirements.
- Broadcom is contributing chip implementation, networking, connectivity, and system-level infrastructure expertise.
- Celestica is assisting with board, rack, and system integration and scalable production systems.
- Data-center partners are expected to host or deploy the infrastructure, but the public announcements do not identify all sites or partners.
This is a broader systems effort, not simply a claim that Broadcom manufactured an OpenAI chip. The companies have not identified the foundry in their public June announcement. Broadcom’s role in networking and connectivity matters because large AI clusters depend on moving data efficiently between processors as well as on the capabilities of each processor. Broadcom’s announcement also describes Celestica’s integration role.
Rank #2
- High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
- Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
- Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
- Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
- Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
What the 10-gigawatt target means
The 10-gigawatt figure describes planned capacity for an accelerator and networking deployment; it is not the electrical draw of one chip and does not establish how many chips will be installed. Converting it into a chip count would require details the announcements do not provide, including chip power, rack design, utilization, cooling, and facility overhead.
It is best understood as an ambitious infrastructure target, not proof of installed compute. Accelerator capacity, a data center’s electrical power, and the compute actually delivered to users are related but different measures.
Rank #3
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
Why OpenAI is pursuing custom silicon
OpenAI and Broadcom say the collaboration is intended to bring OpenAI’s experience with models and products into the hardware. Strategically, custom silicon could give OpenAI more control over how its hardware, software, networking, and serving systems fit together. It could also help diversify accelerator supply and, if the design performs well at scale, improve energy use or serving economics for targeted workloads.
Those are potential advantages, not reported results. A custom chip also brings execution risks: software must support it, production must scale, and the full system must deliver useful work reliably. Efficiency on a benchmark would not by itself establish lower total costs; chip and system prices, software migration, networking, cooling, data-center construction, maintenance, and utilization all affect the economics.
Rank #4
- 48GB AI graphics accelerator
What the performance claims establish—and what they do not
OpenAI and Broadcom say early testing shows substantially better performance per watt than current state-of-the-art hardware. They have not supplied a public benchmark table or independent testing methodology in the cited announcements. Performance-per-watt comparisons depend on the model, precision, batch size, sequence length, memory configuration, networking, cooling, and software stack.
As a result, the claim does not establish that Jalapeño is faster or cheaper than Nvidia across workloads, lowers OpenAI’s total cost, or outperforms other accelerators. Nor does an inference claim establish an advantage in training. The companies also say Jalapeño went from design to production in nine months, but have not detailed which engineering and manufacturing milestones that period covers. It should not be read as confirmation that mass production is complete.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsBest Value
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
What has not been disclosed
The public announcements leave out details needed to independently assess the processor’s capabilities, economics, and production scale. They do not state:
- Process node, die size, or transistor count.
- Memory type, capacity, bandwidth, or supported precision formats.
- Power draw per accelerator, rack density, or network topology.
- Software stack, compiler support, or compatibility with particular models and operators.
- Yield, volume-production status, or the foundry or manufacturing partner.
- Contract value, chip or rack pricing, or confirmed data-center locations.
- Independent benchmark results or whether the architecture supports training as well as inference.
- Whether third parties will be able to buy or rent the processor.
Is Jalapeño a replacement for Nvidia?
There is not enough public evidence to call it an Nvidia replacement. The partnership signals custom infrastructure and possible supply diversification, not an announced exit from external accelerator suppliers. Nvidia GPUs offer flexibility and a mature software ecosystem; a custom ASIC may be more efficient for selected, stable workloads but less adaptable as models and serving patterns change. AMD accelerators, Google TPUs, and AWS chips are other examples of hardware that can coexist in the broader market.
The comparison is not just chip against chip. Software compatibility, networking, memory, system reliability, and the effort required to move workloads matter too. Jalapeño is publicly described as an inference processor, so it should not be assumed to replace the broader training infrastructure required to build frontier models.
Can businesses or developers use Jalapeño?
No public purchase, rental, or developer-access option has been announced. Jalapeño is intended for OpenAI’s infrastructure, and the announcements do not offer a retail card, public cloud instance, sign-up page, or setting that lets ChatGPT users choose the processor. That makes this an infrastructure development, not a consumer chip launch.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchOrganizations evaluating accelerators today need to compare currently available options against their software and deployment requirements. Nvidia GPUs may suit teams prioritizing broad framework support; AMD Instinct is another accelerator stack to evaluate. Google TPU and AWS Trainium or Inferentia are cloud-based options for compatible workloads. None can be compared directly with Jalapeño on price or measured performance from the disclosures available.
Quick Recap
What to watch next
- Whether initial deployment meets the end-of-2026 target and how the broader rollout progresses toward the end of 2029.
- Independent benchmarks that disclose workloads, configurations, and measurement methods.
- Evidence of production volume, system availability, and real-world utilization.
- Published information on software support, pricing, and any access for organizations outside OpenAI.
- Whether the processor remains focused on inference or is later described as supporting training.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




