← AI Business Strategy Daily

AI Innovation

The Next AI Advantage May Be Measured in Watts, Not IQ

OpenAI's first custom inference chip shows why cost, latency, and energy efficiency may shape the next enterprise AI advantage.

By Dr. Anton Gates7 min read10 sources reviewed
A Black female leader in an Eastern University polo discusses AI watts, speed, and cost measures with two business colleagues.

OpenAI named its first custom AI chip Jalapeño. The name is playful. The business signal is not.

On August 25, OpenAI released the first measured performance results for the chip. The company says Jalapeño completed 1.5 to 1.9 times more AI work per watt at peak throughput and produced 1.7 to 3.6 times lower end-to-end latency than the commercial systems used in its comparisons. On highly interactive workloads, OpenAI reported performance gains of 2.1 to 4.1 times.

Those numbers still need to be tested over time, across more workloads, and outside a vendor-led announcement. But they point toward a shift executives should not overlook.

The next phase of AI competition will not be won only by the company with the smartest model. It will also be shaped by who can deliver a useful answer, complete an agentic task, or serve a customer at the best combination of speed, reliability, and cost.

This chip matters more than its name

Inference is what happens after an AI model has been trained. It is the repeated work of processing a request, generating an answer, analyzing a document, completing a coding task, or carrying out the multiple steps required by an AI agent.

Training attracts attention because it produces the next generation of models. Inference is where those models meet customers, employees, and operating budgets.

OpenAI and Broadcom unveiled Jalapeño on June 24 and said initial deployment was planned for the end of 2026. The August 25 results are the first public measurements from the system. OpenAI tested the chip with GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 using InferenceX, a public benchmark maintained by SemiAnalysis.

Jalapeño is rated at 700 watts, although OpenAI says measured sustained power remained at or below 550 watts during the reported tests. The design is optimized for serving language models rather than training them. OpenAI also says it does not plan to sell the chip to other companies because it expects to consume the available capacity itself.

Most executives will never purchase a Jalapeño system. They may still feel its effects.

The real race is the cost of a useful outcome

A chatbot response may require one model call. An agent that researches a customer, checks company policy, updates a system, drafts a response, and verifies its work may require dozens of calls and far more tokens.

That changes the economics quickly.

Gartner estimates that agentic models can consume five to 30 times more tokens per task than a standard generative AI chatbot. The firm expects the unit cost of inference to fall substantially through 2030, while warning that total spending may still rise because demand and task complexity will grow faster.

This is why a faster chip is not merely a data-center efficiency story. Lower latency can make an AI service feel responsive enough for a live customer interaction. Better performance per watt can allow a provider to serve more requests within the same power envelope. Lower cost per task can turn a promising prototype into a service with sustainable margins.

Consider a voice agent that resolves routine insurance questions. If every interaction is slow, expensive, or frequently escalated to an employee, the economics may fail even when the model is impressive. Improve response time, reduce inference cost, and increase completion quality, and the same workflow may become commercially viable.

The competitive question is shifting from “Which model scored highest?” to “Which system produces the best verified outcome at the right speed and cost?”

Do not start shopping for chips

The wrong executive response would be to turn this announcement into an immediate infrastructure project.

OpenAI's results are vendor-reported, based on selected systems and workloads, and Jalapeño has not yet been deployed broadly. The chip also does not remove OpenAI's reliance on Nvidia and other providers. It adds another option to a larger portfolio of computing systems.

For most businesses, the strategic decision is not which accelerator to own. It is how to avoid locking an important workflow to one model, one price structure, or one performance profile.

A procurement team may compare token prices while missing the costs of repeated prompts, long context windows, failed tasks, human review, data movement, and integration. A technology team may choose the fastest model for work that does not require frontier-level reasoning. A product team may celebrate lower unit costs while customer demand causes total consumption to surge.

Dr. Gates's doctoral research emphasizes that digital transformation creates value when leaders connect technology choices with strategy, execution, leadership capability, and customer value. Inference economics makes that connection visible. Technical efficiency matters only when it improves an outcome the organization and its customers value.

Build an inference P&L

Leaders do not need to understand chip architecture to manage this transition. They do need a clearer view of AI unit economics.

Choose one high-volume AI workflow and build a simple operating profile around it. Track:

Then test alternatives. Use the same workload and acceptance criteria across models or providers. Separate tasks that need deep reasoning from tasks that need speed and consistency. Place a cost ceiling on the experiment and review whether lower inference prices are improving margins or simply encouraging more consumption.

This turns AI infrastructure from an abstract technology expense into a management decision.

  • the cost per successfully completed task, not just the cost per token;
  • total response time and the point at which delay harms the user experience;
  • the number of model calls and tokens consumed by a complete workflow;
  • the percentage of tasks completed without rework or human escalation;
  • the business value created, such as revenue, time saved, risk reduced, or service capacity added; and
  • the effect of routing routine work to smaller or less expensive models.

The model leaderboard is no longer enough

Jalapeño will not determine the enterprise AI market by itself. It is evidence of a broader change.

AI providers are integrating models, software, networking, memory, and custom silicon to improve the economics of delivering intelligence. McKinsey identifies custom silicon, model optimization, advanced packaging, and co-packaged optics among the technologies with the greatest potential to reduce inference costs. Those improvements could expand AI from a limited set of high-value use cases into routine operational work.

That expansion will reward businesses that know the economics of their workflows before prices fall.

When AI becomes faster and cheaper, every use case does not automatically become a good use case. The advantage will go to leaders who can distinguish inexpensive activity from valuable execution.

Questions for executives

  1. Do we know the cost per successfully completed AI-enabled task, including rework and human review?
  2. Which workflows would become viable if inference became materially faster or less expensive?
  3. Are we designing AI operations that can route work across models as capability, latency, and pricing change?

Sources and further reading

  1. Jalapeño's first results show industry-leading speed and efficiency in AI inferenceOpenAI · 2026-08-25
  2. The full stack behind abundant intelligenceOpenAI · 2026-08-25
  3. OpenAI and Broadcom unveil LLM-optimized inference chipOpenAI · 2026-06-24
  4. OpenAI and Broadcom Unveil LLM-Optimized Intelligence ProcessorBroadcom Investor Relations · 2026-06-24
  5. Open-Source Agentic Inference BenchmarkSemiAnalysis / InferenceX · 2026-08-26
  6. OpenAI says Jalapeño chip outperforms NvidiaAxios · 2026-08-25
  7. Gartner predicts inference on a one-trillion-parameter LLM will cost providers over 90% less by 2030Gartner · 2026-03-25
  8. The technology shifts reducing AI inference costsMcKinsey & Company · 2026-06-01
  9. Digital Disruption: Exploring the Traits of a Successful Digital Transformation StrategyDr. Anton Gates · 2025
  10. Digital Transformation Strategy: Doctoral Research Poster SessionDr. Anton Gates