OpenAI spent $8.4 billion on servers in 2025. That figure is on track to hit $14 billion in 2026 — the cost of serving inference at ChatGPT scale, paid almost entirely to Nvidia as margin on hardware that no competing AI lab can afford to match at the same volume. The economics were not sustainable, and everyone involved understood that. Jalapeño is the answer to that problem: a custom ASIC built in 16 months with Broadcom, manufactured on TSMC’s N3P node, and presented at Hot Chips 2026 alongside benchmark results claiming 1.5 to 1.9 times more throughput per watt than Nvidia Blackwell. The number is real. The number is also the wrong thing to look at.
The chip is not for sale. You cannot rent time on it. Nobody outside OpenAI runs inference on Jalapeño. It runs OpenAI models for OpenAI products on hardware OpenAI controls. That is the sentence that changes the competitive dynamics, not the benchmark.
What the Chip Actually Is
Jalapeño draws 700 watts per chip at full load. The flagship Nvidia GPU it was benchmarked against draws roughly 1,400 watts. The inference-specific design shows up in everything: MXFP4 precision with a weight-stationary systolic array borrowed from the TPU playbook, HBM4 memory at 15.4 terabytes per second of bandwidth per package, and 32 lanes of 800G SerDes for moving data between chips inside a rack. The full rack configuration ships 128 Jalapeño chips per 2-rack system drawing approximately 160 kilowatts. OpenAI calls the software stack Gluon — built on Triton, with hand-tuned kernels running to about 3,000 lines each — and the internal serving engine Teacup. The design team used an internal version of Codex to generate model-specific kernels — AI assistance reduced the SIMD array area by 8% and the matrix-engine area by 10% — and a chip simulator called chilisim, accurate within 5% of hardware, to iterate without waiting for physical silicon. From hiring to first tapeout was 16 months.
That pace deserves acknowledgment. Google’s first TPU took roughly three years from project start to datacenter deployment. Amazon’s Trainium program ran for years before producing chips that approached parity with Nvidia on key workloads. OpenAI reached A0-stepping silicon in 16 months and posted benchmark results against current-generation hardware within the same year. The chip-design process itself is being accelerated by the models the chips will serve — a concrete instance of AI-assisted engineering that is not hype.
Richard Ho, OpenAI’s VP of Engineering, described the results as showing “a very, very significant performance advancement.” The SemiAnalysis team, who ran the benchmarks, described Jalapeño as part of a “multigenerational platform integrating AI products, models, chips, and memory” — framing that came directly from OpenAI. That framing is the signal most coverage missed.

The Honest Read on the Benchmark
SemiAnalysis ran InferenceX with OpenAI engineers in OpenAI’s lab, on hardware OpenAI configured, using numbers OpenAI provided. The headline numbers from the InferenceX run extend well beyond throughput: end-to-end latency ran 1.7 to 3.6 times lower than the Blackwell baseline, and interactive chatbot workloads came in 2.1 to 4.1 times faster. The Register opened its coverage with “take all of these claims with a grain of salt.” The top comment on the Hacker News thread — which reached 493 points and 316 comments — called the SemiAnalysis piece “a lot more like an OpenAI press release” and noted the account of how the benchmark was run shifted midway through the article. The benchmark is not fabricated, but it was not independently run. Those are different things.
The fairer comparison is also not Blackwell. Jalapeño uses HBM4 memory. Nvidia’s Vera Rubin accelerator uses HBM4 memory. Those are the correct peers. Against Rubin, the total cost of ownership per token comes out roughly even — SemiAnalysis’s own conclusion, which received far less coverage than the Blackwell comparison. Jalapeño beats Rubin on some single-turn throughput metrics even though Jalapeño has not yet implemented multi-token prediction or speculative decoding, both of which Nvidia’s systems use. That’s a meaningful architectural data point: the first-generation chip, without optimizations the GPU competition already ships, is already at TCO parity with Nvidia’s current flagship. But parity is not a blowout, and the Blackwell framing in most headlines was not precise.
There are also categories of workload the benchmark does not cover. AgentX — SemiAnalysis’s multi-turn, long-context production-representative benchmark — was not completed for Jalapeño. The models tested were smaller than current frontier open models like DeepSeek V4 Pro and Kimi K3. The competing GPU systems in the comparison deliver 1.46x to 2x more raw compute and up to 12% more memory than Jalapeño; the efficiency advantage is real, but peak compute headroom is not. Vera Rubin is already shipping to customers. Jalapeño is engineering samples. Production deployment starts late 2026 in small volumes, with broader rollout through 2027.
None of that makes the chip less significant. It makes the benchmark a narrower claim than the coverage suggests. What’s significant is the combination: a first-generation chip at TCO parity with Nvidia’s newest accelerator, designed in 16 months, by a company that had never made silicon before.

What Changes for Everyone Not Named OpenAI
Nvidia’s high-end accelerators run at gross margins that analysts have estimated around 75% for AI hardware. That margin is the cost that every lab buying Nvidia GPUs for inference is transferring to Santa Clara on every query their users run. The analysis is straightforward: a large share of what a lab spends on inference leaks out as supplier margin. The claimed 50% cost reduction from Jalapeño, if it holds at scale, means OpenAI’s per-inference cost falls by half while competitors’ per-inference costs stay where they are. That gap compounds. Every quarter that OpenAI runs on proprietary silicon, every model update that can be deployed more cheaply, every pricing move that competitors cannot match — all of it follows from this structural position.
Google has had TPUs since 2016. Amazon has Trainium and Inferentia. Microsoft co-invests with OpenAI and shares access to the same hardware roadmap. Anthropic, Mistral, and the rest of the AI lab ecosystem are renting Nvidia at rack scale. The competitive picture is not OpenAI versus Nvidia; Nvidia still makes the GPUs OpenAI uses for training, and OpenAI’s CFO has explicitly framed Jalapeño as complementary to existing vendor relationships. The picture is OpenAI versus the labs that don’t own their inference substrate — and that is most of them.
There is a version of this story where custom silicon doesn’t matter because the models change faster than the hardware. If a new architecture makes a chip obsolete within a single chip generation, capital invested in proprietary ASIC development is wasted. OpenAI’s own roadmap is an argument against that reading: they describe Jalapeño as the first generation of a multigenerational platform. The B0 stepping of the chip, with a claimed 25% improvement in performance per watt over A0, is already in fabrication. The software stack — Gluon, Teacup, chilisim — is being built for longevity, not for one chip cycle. The loop the analysis describes is plausible: cheaper inference funds more training experiments, which produce better models, which inform better chip design. It is not yet demonstrated at scale, but it is what Google and Amazon claim their TPU and Trainium programs produce.
The counterargument worth taking seriously is about manufacturing, not design. Designing a chip is hard. Getting it manufactured at high yield, at scale, with the supply chain reliability to actually displace GPU procurement — that is a different problem, and it is where a number of previous custom-silicon programs have stalled. Amazon’s early Trainium chips underperformed targets badly enough that the company ran on Nvidia for years while fixing the stack. OpenAI taped out Jalapeño in November 2025 and is presenting A0 stepping results nine months later. The B0 production ramp in 2027 will be the first real test of whether the program survives contact with volume manufacturing. The Decoder noted that Vera Rubin is already in customer racks; OpenAI is not. The gap between engineering samples and datacenter-grade production is where the real execution risk lives.
The Number That Will Matter More
OpenAI’s first-generation Jalapeño chip is not the story. The second one is. If B0 ships in 2027 with the claimed efficiency gains, if the production ramp works, if speculative decoding and multi-token prediction get implemented on the Gluon stack — and if the third chip follows on a similar cadence — OpenAI will have built something that took Google a decade to make meaningfully competitive with Nvidia. That timeline compression is the real achievement to watch, not the benchmark against Blackwell.
The structural consequence is already in motion. OpenAI is now a company that designs hardware, writes custom kernels, builds serving infrastructure, and deploys all of it at a scale that no competitor can access. The labs that will matter in 2028 are the ones that figure out an answer to a cost structure that just changed — either through their own silicon programs, through deep relationships with cloud providers who have custom hardware, or through model architectures efficient enough to stay competitive on commodity GPUs. None of those paths are easy. OpenAI just made them necessary.

AI-generated editorial illustration · TemperatureZero · August 26, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive