OpenAI stated that its self-developed AI chip, Jalapeno, surpasses NVIDIA's GB300 in power efficiency and response latency. Designed specifically for inference workloads and developed with assistance from Broadcom, the chip delivers high performance at a low power consumption of 700 watts, aiming to reduce computing costs. OpenAI noted that the chip offers both high throughput and low latency, making it particularly suitable for large models. However, OpenAI has no intention of completely replacing NVIDIA and will continue to maintain a multi-vendor strategy.
OpenAI is accelerating the development of its proprietary AI chips, with its Jalapeno chip outperforming NVIDIA's GB300 in key benchmarks and expected to enter practical use as early as later this year.
According to Bloomberg on August 25, OpenAI stated that Jalapeno outperforms NVIDIA's GB300 in two key metrics: AI workload processing per unit of power consumption and response speed. The chip was jointly developed by OpenAI and $Broadcom (AVGO.US)$ Broadcom, primarily to support AI model inference, and is expected to help OpenAI reduce its reliance on external chip suppliers.
Unlike traditional chips that often require a trade-off between throughput and latency, Jalapeno delivers both high throughput and low latency. Richard Ho, OpenAI’s head of chip development, noted that the chip achieves robust performance at a low power consumption of 700 watts, helping to reduce electricity costs for data centers.
However, OpenAI does not position Jalapeno as a comprehensive replacement for NVIDIA chips. Richard Ho stated that OpenAI’s demand for computing power is immense, and the company will continue to collaborate with multiple suppliers in the foreseeable future, with NVIDIA remaining a key partner.
Meanwhile, the second-generation Jalapeno has reached a mature stage of development, and conceptual design for the third-generation chip has already begun. OpenAI is seeking to further reduce AI infrastructure costs through its in-house chip development efforts.
Benchmark Performance: Outperforms NVIDIA GB300 in Two Key Metrics
According to evaluations conducted by OpenAI using public benchmarking systems, Jalapeno outperforms NVIDIA's GB300—the highest-ranked product in the testing system—in both power efficiency and response latency.
In an interview, Richard Ho stated that Jalapeno excels in high-throughput scenarios, enabling service to more users at lower costs. Additionally, its strong low-latency performance can significantly reduce response times for customers who prioritize speed.
OpenAI conducted public tests on the chip using one of its small open-source models, as well as third-party models from DeepSeek and Moonshot AI. Jalapeno showed the most pronounced advantage when running Moonshot AI’s Kimi model, which was the largest model included in the tests. Internal testing also indicated that the chip performs well when running several large, advanced models developed by OpenAI that have not yet been released.
In a blog post, OpenAI noted that this demonstrates the chip’s design "becomes increasingly valuable as workloads grow in size and complexity." Notably, NVIDIA’s latest generation chip, Vera Rubin, which has just begun shipping, was not included in these tests.
Positioned for the inference stage; does not replace existing suppliers.
Jalapeno is specifically designed for the AI inference phase—the stage where models respond to user instructions and execute tasks after training is complete. It is not intended for AI model training, which remains a core strength of NVIDIA’s technology.
In the inference domain, Jalapeno’s speed advantage enables OpenAI to achieve performance levels that previously required specialized memory architectures. While OpenAI currently uses technology from Cerebras Systems for some models, such chips are better suited for smaller-scale models.
Richard Ho stated that Jalapeno can handle larger-scale models, thereby creating a differentiated competitive advantage.
Nevertheless, Richard Ho explicitly clarified that OpenAI will not abandon its existing suppliers in the near term. He noted that the company’s demand for computing power is immense, which is why it has contracted with numerous diverse suppliers—a situation expected to persist for some time.
He also emphasized that NVIDIA remains a key partner: “NVIDIA is an excellent partner, and we continue to require substantial volumes of NVIDIA products.”
R&D acceleration: Second generation already in development
OpenAI stated that it has further accelerated the R&D process for Jalapeno by leveraging its own AI models. The collaboration between Jalapeno and Broadcom was announced last year, with both companies highlighting the chip’s record-breaking development speed in June of this year.
Currently, the second-generation chip has reached a relatively mature stage. Richard Ho indicated that tape-out is expected to be completed within the “next few months.” Meanwhile, conceptual design work for the third-generation chip has already begun, aiming to further reduce the costs associated with OpenAI’s large-scale global deployment of AI infrastructure.
Richard Ho stated that the company has achieved power consumption and cost levels that can tangibly reduce infrastructure expenses, noting that “this is only the first step.”
Editor/lambor