OpenAI has moved beyond reported plans for custom silicon. On June 24, 2026, it and Broadcom unveiled Jalapeño, an OpenAI-designed processor aimed primarily at large-language-model (LLM) inference. Engineering samples are running in OpenAI laboratories at their target frequency and power, and initial deployment is planned by the end of 2026. The chip is part of a much larger, 10-gigawatt infrastructure program announced in October 2025—not a retail product that developers can buy today.
The short version
- Processor: Jalapeño, OpenAI’s first named “Intelligence Processor.”
- Main job: Serving trained LLMs for products such as ChatGPT, Codex and API applications.
- Architecture: OpenAI designs the accelerator and system; Broadcom provides silicon implementation, networking and connectivity expertise.
- Systems: Celestica is supporting board, rack and server-system integration.
- Manufacturing: Reuters reported that TSMC will manufacture the chips.
- Scale: A forward-looking plan for 10 gigawatts of accelerator and networking infrastructure, deployed from the second half of 2026 through the end of 2029.
- Availability: No public price, developer kit, customer order page or third-party rental offering has been announced.
OpenAI’s objective is to reduce the cost, latency, power use and supply risk of operating its models at very large scale. That makes Jalapeño a diversification effort alongside Nvidia, AMD and cloud-provider hardware, not proof that OpenAI has replaced the GPU ecosystem.
OpenAI’s June 24, 2026 announcement describes Jalapeño as the first generation of a broader, multi-generation platform.
What OpenAI and Broadcom announced
The October 2025 infrastructure partnership
On October 13, 2025, the companies announced a collaboration to deploy 10 gigawatts of OpenAI-designed accelerators and Broadcom networking systems. The racks are intended for OpenAI facilities and partner data centers. The stated schedule begins in the second half of 2026 and targets completion by the end of 2029.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
OpenAI is responsible for accelerator and system architecture. Broadcom contributes implementation expertise plus Ethernet, PCIe, optical and other connectivity technologies. The companies did not disclose financial terms or a detailed facility-by-facility procurement schedule. Read the October 2025 partnership announcement for the companies’ stated scope and timetable.
The June 2026 Jalapeño unveiling
Jalapeño is the first publicly named processor from that program. OpenAI says engineering samples are operating in its labs at production target frequency and power, including workloads from GPT‑5.3‑Codex‑Spark. Initial deployment is targeted for the end of 2026. The company says the design-to-tape-out cycle took nine months and that its models helped accelerate parts of chip design and optimization.
Inference, not a general-purpose replacement
Inference is the stage where a trained model generates an answer, code completion, image, prediction or other output for a user or application. Training is the resource-intensive process of creating or updating the model. OpenAI’s public description of Jalapeño focuses on interactive inference for ChatGPT, Codex, APIs and future agentic products.
The processor is designed around the kernels, memory movement, networking, scheduling and serving behavior that dominate those workloads. OpenAI says it is built from scratch for modern LLM inference and is intended to support current and future LLMs across the industry. That does not establish that it is equally suitable for model training, scientific computing, graphics or unrelated AI workloads.
Who is doing what?
| Participant | Confirmed role |
|---|---|
| OpenAI | Accelerator and system architecture; model, kernel, serving and product expertise |
| Broadcom | Silicon implementation, networking, connectivity and large-scale deployment expertise |
| Celestica | Board, rack, server-system integration and scalable production systems |
| TSMC | Chip manufacturing, according to Reuters reporting |
| Microsoft and other data-center partners | Referenced by Broadcom as participants in gigawatt-scale deployment; a complete allocation has not been disclosed |
OpenAI is therefore designing the chip, not operating an independent semiconductor factory. Implementation, manufacturing, packaging, systems and data-center deployment still depend on outside partners.
What 10 gigawatts does—and does not—mean
Ten gigawatts describes the planned electrical scale of many accelerator racks and their networking systems. It is not the rating of one chip, one server or one already-operating facility. The figure is a forward-looking deployment target.
Rank #2
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
- It does not reveal the final chip count, memory capacity, tokens per second or cost per million tokens.
- It does not prove that 10 GW is installed or operational today.
- It does not disclose total project cost, facility locations or the exact construction and procurement schedule.
The announced window runs from the second half of 2026 to the end of 2029, so the infrastructure ambition should not be confused with delivered compute capacity.
Why OpenAI wants custom silicon
Lower serving cost and better utilization
A specialized accelerator can omit general-purpose functions that OpenAI does not need for a known, high-volume workload. At sufficient scale, even modest gains in utilization can materially affect operating costs.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Performance per watt and latency
OpenAI and Broadcom say early testing shows substantially better performance per watt than the current state of the art. Final figures and the test methodology have not yet been published. Lower power use and faster response times are especially valuable for conversational, coding and agentic applications.
Supply diversification
Owning more of the design gives OpenAI another path when advanced accelerators, memory, packaging or data-center capacity are constrained. It does not remove the company’s dependence on Nvidia, AMD, cloud providers, Broadcom, TSMC, memory suppliers or facility operators.
Hardware-software co-design
OpenAI can tune the processor to the models, kernels and serving stack it operates instead of adapting every workload to a standard accelerator. That tighter control is the central technical rationale for the project.
Is Jalapeño better than Nvidia?
There is no independent, reproducible benchmark in the cited announcements that settles that question. Broadcom CEO Hock Tan told Reuters that the chip is as good as Nvidia Blackwell and Google TPUs; that is an executive comparison, not a published third-party result. OpenAI’s efficiency statement is likewise based on early internal testing.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #3
- Built for Running LLMs Locally: RDNA 4, 128 AI Accelerators, up to 1,531 TOPS (INT4) for fast inference and fine-tuning
- 32GB GDDR6 VRAM for Large AI Models: 256-bit, up to 640GB/s bandwidth, run large language and multi-modal AI models without offloading
- Multi-GPU Scaling for Local AI Clusters: PCIe 5.0 and 2-slot design support dense multi-GPU builds for local AI training and inference clusters
- Diecast Shroud and Backplate: Wave-pattern design cuts memory temperature by up to 16%, keeping clocks steady during long AI training runs
- Phase-Change GPU Thermal Pad: Delivers superior thermal conductivity for consistent performance and longevity under heavy AI loads
The eventual comparison must specify model, precision, batch size, latency target, software stack, power boundary and whether the test measures a single accelerator or a complete rack. A chip optimized for OpenAI-style LLM inference could excel there while being less useful for training or non-LLM workloads. Until the promised technical report and independent measurements appear, claims that Jalapeño is the “fastest AI chip” or has defeated Nvidia are premature.
Will it replace Nvidia GPUs?
Probably not in the near term. Nvidia’s advantage includes CUDA, libraries, developer familiarity, broad workload support, manufacturing scale and a large installed base. OpenAI’s announcement presents Jalapeño as an addition to its infrastructure strategy and says the company will continue working with ecosystem partners.
The more defensible interpretation is that OpenAI is trying to reduce its dependence on off-the-shelf accelerators for selected, high-volume inference workloads. Nvidia, AMD and cloud GPUs are likely to remain important for training, experimentation, heterogeneous applications and capacity that Jalapeño does not cover.
What remains unknown about the processor
The public releases do not state:
- Process node, die size or transistor count
- Memory type, capacity or bandwidth
- Exact power draw, latency, throughput or tokens per second
- Throughput per rack or cost per million tokens
- Production volume, yield or reliability results
- Whether OpenAI will sell the chip or offer it through an outside cloud
- How it performs against Nvidia Blackwell, Google TPU, AMD Instinct or other accelerators under a common test
Tape-out is an important design milestone, but it is not the same as high-volume production, qualification or successful data-center deployment.
Free tools Windows power users keep installed
One-click scans. No signup required.
Who will be able to use Jalapeño?
Current evidence points to OpenAI’s own infrastructure. Reuters reported that the chips and server systems will be used only by OpenAI. OpenAI’s statement that the architecture is flexible enough for current and future LLMs across the industry describes design intent, not a promise of commercial access.
| Access question | What is established |
|---|---|
| Can a developer buy a processor? | No public purchase channel announced |
| Can a cloud customer rent Jalapeño? | No cited provider has announced availability |
| Will OpenAI sell systems? | Not announced |
| Is it consumer hardware? | No evidence |
What it could mean for ChatGPT and the API
If production deployment meets the companies’ goals, possible effects include lower serving costs, more capacity during demand spikes, lower data-center power requirements and lower latency for interactive workloads. Those are potential operational outcomes, not announced price cuts or guaranteed speed improvements.
Rank #4
- 24GB GDDR7 ECC Memory: handles large AI, 3D and rendering files smoothly
- Powerful CUDA Compute - 8,960 CUDA cores for fast graphics and computing power
- AI & Ray Tracing Boost - Tensor of the 5th generation and RT cores of the 4th generation
- PCIe 5.0 x16 interface - fast data connection with modern systems
- 4 × DisplayPort 2.1 - Multi-monitor support for professional workflows
For customers, the near-term buying decision remains between managed access—such as the OpenAI API or ChatGPT Business/Enterprise—and operating or renting conventional accelerators. Jalapeño-specific pricing has not been disclosed.
Timeline
- 2023: Reuters reported that OpenAI was exploring its own chip effort.
- Early 2024 (reported): Work with Broadcom reportedly began roughly 18 months before the October 2025 announcement, according to comments cited by AP.
- October 13, 2025: OpenAI and Broadcom announced the 10-GW collaboration.
- June 24, 2026: Jalapeño was unveiled, with engineering samples running in OpenAI labs.
- End of 2026: Initial deployment target.
- End of 2029: Target for completing the broader deployment program.
The early-history details are reported rather than fully specified in OpenAI’s public chip announcement; the partnership dates and deployment targets come from the companies’ releases.
Recommended Free Tools
Commercial implications for infrastructure buyers
Jalapeño is not currently a hardware purchasing option. Buyers choosing infrastructure today should evaluate established alternatives according to workload and control requirements:
- Nvidia data-center GPUs for broad training and inference compatibility and the mature CUDA ecosystem.
- AMD Instinct for a non-Nvidia accelerator path, with software-porting requirements assessed in advance.
- AWS AI infrastructure, Google Cloud TPU or Microsoft Azure ND-series infrastructure when renting capacity is preferable to owning systems.
Cloud prices vary by region, instance, contract and availability. None of those pages establishes that Jalapeño is offered through the service.
Bottom line
OpenAI’s Broadcom partnership has produced a real, named inference processor and moved the company further down the infrastructure stack. Jalapeño could improve OpenAI’s cost, power and capacity economics if it reaches production at the claimed efficiency. The strategic significance is clear; the competitive verdict is not. That verdict depends on production-scale performance, software support, supply execution and cost data that have not yet been published.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.




