Skip to content
OCopenai.com·

The Work Now Within Reach

AI summary

OpenAI's new custom inference chip, Jalapeño, significantly enhances AI capabilities by offering 1.5 to 1.9 times more peak token throughput per watt and 1.7 to 3.6 times lower end-to-end latency compared to commercial systems in InferenceX tests. This development, alongside accelerators from NVIDIA, AMD, and other partners, aims to expand what people and businesses can achieve with AI, making previously impractical ideas feasible and opening doors to new breakthroughs and possibilities. Deployment is planned to begin by year-end.

Why this one

This report is the first time OpenAI has announced its own custom inference chip, Jalapeño, which outperforms commercial systems in throughput and latency.

Time & source

Published
09/08, 13:00 UTC+0
Ingested
09/09, 21:00 UTC+0
Source type
Official
Tier
First-party
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Article

AI is expanding what people and businesses can achieve. Ideas once held back by time, cost, or access to expertise are becoming practical to pursue. That opens the door to breakthroughs in human achievement, to new businesses, new markets, and new possibilities. OpenAI is built to help realize these ambitions.

GPT‑6 Astra ⁠ is a major step forward in AI capability. It is the world’s most intelligent and aligned model, and is state-of-the-art in areas such as computer use, browsing, software engineering, cybersecurity, science, and professional work.

OpenAI is built to turn this research progress into broad benefit for people and organizations around the world. Our broad reach across both consumer and enterprise means new capabilities have a direct route to customers through products they already use. And our full stack compute strategy gives us greater control over the capacity, performance, and the cost of delivering that intelligence.

These advantages compound. Better models open up new work. More efficient compute makes that work affordable at a greater scale. Revenue from growing adoption funds further research and infrastructure, strengthening our ability to pursue the next breakthrough and bring it to people everywhere.

Consumer and enterprise strengthen each other

Our products reach more than one billion weekly active users ⁠ and 2.5 million businesses. Every research advance improves ChatGPT, ChatGPT Work, Codex, and applications built on our API. One investment in model research supports many products and their diversified revenue streams.

People who know ChatGPT at home bring that familiarity to work. Enterprise deployments give them tools suited to complex work and the requirements of their organization. Experience at work can then change what they expect from AI in their personal lives. Developers extend our reach further by building applications for needs we would not have identified ourselves. And over time we expect there to be a continuous blurring between these traditionally more distinct segments as our agentic products get to know you better as a person but at work and in your personal life.

Individual use and enterprise deployments expand over time. In our study of people on individual ChatGPT plans ⁠, daily message volume was roughly 50% higher six months after signup than in the first month, and people had tried roughly twice as many distinct tasks.

Indexed to the first 28 days after signup

Source: OpenAI

Our business model lets us earn revenue as that use grows. Free access, supported by advertising, helps people discover where AI is useful. Subscriptions and usage-based offerings let customers spend more as they find more value.

More work becomes worth doing

Increasing model capability expands what customers can afford to do. A business can fix a process that once demanded too much specialist time or enter a market it could not previously serve profitably.

We see this in the work of our own team. A recent update from our research team ⁠ shows how researchers are contributing code faster and running more experiments while delegating increasingly complex tasks to agents. Teams are leveraging agents to resolve infrastructure problems that once required specialist support. The research organization now uses 3.1 agent-workdays of effort for every workday of human labor.

Source: OpenAI Research Organization, mid-August 2026, Stanford 8-hour workday.

People still set research priorities and judge results, but agents give them more capacity to pursue promising ideas. For OpenAI, that work can help improve the models and the infrastructure that our customers depend on.

The possibilities extend to scientific discovery. We recently announced that one of our internal models has produced a solution to the Navier–Stokes Millennium Prize Problem ⁠, one of mathematics’ deepest open questions that has remained unresolved for roughly 90 years. This marks a significant milestone in AI’s ability to contribute to mathematical research, and shows how we can empower scientists to advance research and technology.

As AI expands human capability, it opens a market far beyond the work we do today. More people can pursue their biggest ambitions, make new discoveries, and build things that improve lives for generations to come.

Customer stories

1 of 5

Compute that delivers better value

Serving expanding demand requires both capacity and better economics. Our compute strategy ⁠ spans data centers, chips, software, models, and products. We manage those choices together, continually seeking the strongest combination of capability, speed, reliability, efficiency, and cost for each workload.

Training a frontier model and running a fast, interactive agent place different demands on infrastructure. We build where designing components together improves performance or cost, and partner where others offer the strongest option. This gives us the flexibility to choose the best system for each task.

GPT‑5.6 Sol helped improve our production serving software, reducing end-to-end serving costs by 20% ⁠. Additional improvements increased token-generation efficiency by more than 15 percent, allowing the system to produce more output from its compute.

Source·openai.com