Categories
Technology

OpenAI’s new Chip promises faster AI

Custom processor promises faster responses, lower power use and wider deployment from 2027

OpenAI is taking a bigger role in the hardware powering artificial intelligence after revealing the first performance results of its custom AI chip, Jalapeño. The company says the processor can deliver more AI work while using less power and can significantly cut the time users wait for responses.

The results are important because OpenAI is one of the biggest users of advanced computing hardware, particularly for running AI models at scale. Instead of depending entirely on chips supplied by other companies, OpenAI is now developing its own silicon specifically for AI inference — the process of taking a trained model and using it to respond to user requests.

OpenAI developed Jalapeño with semiconductor and networking company Broadcom. The chip is not being positioned as a general-purpose processor. It has been designed around the specific demands of large language models and AI applications, where speed, memory movement and power consumption can have a major impact on operating costs.

The company tested Jalapeño using InferenceX, a public benchmark developed by SemiAnalysis to measure AI inference performance. OpenAI tested the chip with three large models: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T.

The results were striking. OpenAI said Jalapeño delivered between 1.5 and 1.9 times more AI work per watt at peak throughput compared with the Nvidia systems used in the tests. It also recorded between 1.7 and 3.6 times lower end-to-end latency. In simple terms, the chip was able to process AI workloads more efficiently while reducing the time taken to produce responses.

That combination matters because AI companies face two major challenges as usage grows: computing capacity and electricity consumption. A system that can produce more output using less power can potentially serve more users without requiring a proportional increase in infrastructure.

OpenAI said Jalapeño has a 700-watt rating, while its measured sustained power remained at or below 550 watts during the workloads tested. The company compared its performance with Nvidia’s high-end GB200 and GB300 systems. On the three tested models, OpenAI reported substantial improvements in performance per watt and response latency.

The company is particularly interested in reducing latency because AI is increasingly being used for interactive tasks. Chatbots, coding assistants and AI agents need to respond quickly, especially when they are carrying out multiple steps on behalf of a user. A delay of a few seconds may become much more noticeable when an AI agent has to make several calls before completing a task.

OpenAI’s approach is to design the entire system around AI inference rather than treating the chip as an isolated component. The company has worked on the processor, memory, networking and software as a combined system. This allows engineers to address bottlenecks that can occur when information has to constantly move between different parts of an AI data centre.

Memory is particularly important for modern AI models. During inference, large amounts of information must be moved and accessed quickly as the model generates an answer. OpenAI has therefore designed Jalapeño to keep critical data closer to where it is required, reducing some of the communication overhead that can slow down AI systems.

The company also says artificial intelligence itself played a role in developing the chip. OpenAI engineers used its AI tools to explore designs, improve software and speed up verification. According to the company, Jalapeño moved from initial design to tapeout in about nine months.

OpenAI also used Codex and GPT-Astra to optimise software for several open-weight AI models. The company reported that some AI-generated implementations for specific model components were 1.5 to 1.8 times faster than existing implementations created by human engineers. However, these figures apply to selected components rather than complete AI models, so they should not be interpreted as an overall model-speed increase.

Despite the strong benchmark numbers, Jalapeño is not expected to replace Nvidia hardware across OpenAI’s infrastructure. The company has made it clear that Nvidia accelerators and chips from other suppliers will continue to be used. Jalapeño is instead being developed as another option that gives OpenAI greater control over how its AI services are powered.

That distinction is important because the current results are based on specific benchmark conditions and selected models. Actual performance at large scale will depend on factors including software optimisation, networking, workload patterns and how the chips perform inside production data centres.

OpenAI plans to begin deploying Jalapeño within its infrastructure by the end of 2026, while wider deployment is expected in 2027. The company is already working on future generations, indicating that Jalapeño is intended to become a long-term hardware platform rather than a one-off experiment.

The move also reflects a broader shift across the AI industry. Technology companies are increasingly developing custom AI chips to improve performance, manage energy use and reduce dependence on a small number of hardware suppliers.

The Jalapeño project is ultimately about having more control over the infrastructure behind ChatGPT and its other AI products. If the early performance gains translate into real-world production workloads, custom silicon could help the company deliver faster AI responses while keeping the enormous cost of running AI systems under greater control.

Jalapeño remains an internal accelerator rather than a chip OpenAI plans to sell commercially. But with wider deployment planned for 2027 and newer generations already being developed, its progress could become an important part of the race to build faster and more efficient AI infrastructure.

 

Leave a Reply

Your email address will not be published. Required fields are marked *