Blog

GPT-6 Astra: What Really Changes for Your Business

OpenAI launched GPT-6 on September 3. The same benchmark scored 62.7% and 99.9% depending on the harness, and token price rose 2.5x.

September 04, 2026 · Agência Primeira Página

GPT-6 Astra: What Really Changes for Your Business

On September 3, 2026, OpenAI launched GPT-6 Astra, which it describes as "the smartest and most aligned model in the world." It started with a limited group of organizations and, according to the announcement, will reach all ChatGPT Plus, Pro, Business, and Enterprise plans in the coming days, as well as the OpenAI API, Microsoft Azure, and AWS Bedrock.

The company's president, Greg Brockman, told journalists that "it's not unreasonable to feel that we are now in the era of AGI." That's the line that made headlines. Beneath it are three things that change the decision for anyone using AI to work: what the machine can now do, what the numbers actually mean, and how much it costs.

What It Does That the Previous Model Didn't

The leap OpenAI highlights is computer use: filling out an on-screen form, updating a customer record in the CRM, organizing a calendar, researching and drafting a summary inside a text editor, building a website and then checking that the buttons work. It's the model operating the interface, not just returning text for you to execute.

The numbers it published, all compared with GPT-5.6 Sol, the previous generation:

  • OSWorld 2.0 (real-world computer tasks): 72.6% versus 65.7%, in about 47% less time per task — roughly 40 minutes versus 75.
  • ScreenSpot-Pro (finding and clicking the right element in a cluttered interface): 92.7% versus 76.9%.
  • Agents' Last Exam (professional work in real software): 59.3%, versus 55.5% for Claude Opus 5 and 53.6% for Sol.

Technical specs: a context window of 1,050,000 tokens, output of up to 128,000 tokens, and a knowledge cutoff of April 30, 2026.

The Same Benchmark, Two Numbers: 62.7% and 99.9%

In the announcement, OpenAI writes that Astra "saturates ARC-AGI-3 with 99.9%." ARC-AGI-3 is a reasoning benchmark set in novel environments, maintained by the ARC Prize Foundation — an organization outside OpenAI. And ARC Prize published, on the same day, two results for the same model:

  • 62.7% on the standard harness, the minimal interface that's the same for every model, at a cost of $26,098 in calls.
  • 99.9% with a "provider adapter," a vendor-specific harness that preserves reasoning state between calls and compresses the conversation, letting the model reuse work already done. It cost $18,817.

This isn't fraud, and it isn't a minor detail. Across the 167 matches both harnesses solved, the adapter version was about 3.66 times faster and used 49% fewer tokens. And the standard harness's 62.7% is, on its own, already the best result ARC Prize has measured: Claude Opus 5 scored 30.2% and GPT-5.6 Sol, 7.8%. ARC Prize itself writes that saturating this benchmark wouldn't be proof of AGI, because its scope is narrow compared with the real world. To add to the confusion, part of the press reported 98.6%, which is yet another configuration of the same measurement.

The practical lesson applies to any vendor. An AI benchmark number isn't a property of the model: it's the result of the model plus the scaffolding it runs on. When someone shows you a percentage, the two questions that separate information from marketing are "measured with which harness?" and "at what effort level?" We wrote about this kind of trap in how to evaluate an AI.

The Price Jumped a Tier, and the Math Isn't Per Token

Via the API, Astra costs $10 per million input tokens and $50 per million output tokens, with cached input at $1. Fast mode costs double, Batch and Flex cost half, and above 272,000 input tokens the math changes again. That's about 2.5 times the price of GPT-5.6 Sol.

Brockman responded to the price hike by saying that "what you actually want is price per task." The argument is legitimate: if the task finishes with fewer tokens and in less time, total cost drops even with a pricier token. OpenAI shows a case like this in Terminal-Bench Science, where Astra scores 64.6% versus 52.6% for Claude Fable 5.1, with an estimated API cost about 31% lower.

But Artificial Analysis, which measures independently, reached the opposite result on its overall index: Astra uses fewer tokens than Sol for similar performance, yet still comes out more expensive, because the price increase outweighs the savings. On that firm's intelligence index, Astra scores 61 versus 66 for Claude Fable 5.1.

In other words: "price per task" is the right criterion, and that's exactly why the vendor can't be the one to answer it. Price per task is measured on your own task. If you already pay per token in production, this launch is a reason to redo the spreadsheet, not to switch models by default. The cost reasoning is detailed in AI prices dropped 80%.

First Time at the "Critical" Security Level

Astra is the first OpenAI model to reach the Critical level in cyber capability under the company's own Preparedness Framework. Tested without production safeguards, it scored 100% on ExploitBench (versus 78.5% for Sol) and 42.4% on ExploitGym (versus 30.3%). On SRE-Bench, which measures binary reverse engineering, it solved 88.0% of tasks on the first attempt and 99.2% within four attempts — versus 55.9% and 68.7% for Sol.

The most concrete data point: in an evaluation built from flaws found in the previous three months, the model discovered and exploited two previously unknown zero-day vulnerabilities. OpenAI says it is disclosing both to the maintainers of the affected software.

The version being released refuses offensive tasks — creating a proof-of-concept exploit, for example — and accepts defensive work, such as code review and patch application. Through the Daybreak program, the company plans to loosen these restrictions for defense teams in the coming weeks.

For an ordinary business, the effect isn't abstract: finding a flaw got much cheaper, while fixing one still depends on people and a maintenance queue. This is the gap we had already measured in AI finds flaws faster than companies fix them.

What OpenAI Didn't Put in the Headline

  • The model got harder to monitor. The company itself writes that Astra's written reasoning is harder to monitor than Sol's, because it controls what it writes more tightly and solves problems with fewer visible steps. It classifies this as a regression it takes seriously.
  • Not every chart is favorable. On Humanity's Last Exam with tools, Astra scores 57.2% versus 65.0% for Claude Fable 5.1. In agentic coding on DeepSWE, Meta's Muse Spark 1.3 comes out ahead with 75.4% versus 74.1%.
  • One of the headline benchmarks has a history of conflict. FrontierMath, where Astra scores 98%, was created by Epoch AI with funding from OpenAI — an arrangement disclosed only at the end of 2024 that drew public criticism from mathematicians who had collaborated without knowing about it.
  • Almost every number came from maximum effort, a setting that boosts the score and isn't the one you use day to day.

None of this undermines the launch. It's a reminder that whoever publishes most of the numbers is an interested party — and that independent verification, when it exists, tends to arrive later.

What to Do This Week

  1. Don't switch models by default. Pick the ten tasks your team repeats most and run the same ten on the current model and the new one, measuring cost per completed task and how many came back for rework.
  2. Revisit your budget if you pay per token. A 2.5x jump in list price changes your projections, even if consumption drops.
  3. Start with screen-based work. That's where the measured gain is largest: filling out forms, updating records, checking spreadsheets, verifying that a website still works after a change. Pick a boring, repetitive task you already know how to time.
  4. Treat security as a queue, not as news. If flaw discovery sped up, the bottleneck becomes your update and patching routine.
  5. Require logging. An agent that operates a system needs to leave a trail of what it did, with human approval on actions that spend money or delete data.

This is the kind of approach — measured tasks, cost per task, human approval in the right place — that we apply in AI implementation for businesses. The useful question isn't which model is the best in the world, but which task in your business is worth handing to a machine this week.

Sources

Official announcement and security page for GPT-6 Astra (OpenAI, September 3, 2026); OpenAI's model and API pricing documentation; ARC-AGI-3 analysis published by the ARC Prize Foundation; independent measurements from Artificial Analysis; coverage from Fortune, Engadget, and The New Stack; commentary from Gary Marcus; and the FrontierMath funding history published by Epoch AI.

Frequently asked questions

What is GPT-6 Astra and when was it released?

It is OpenAI’s new frontier model, announced on 3 September 2026 as “the world’s most intelligent and aligned model”. It started with a limited set of organisations and, according to the announcement, reaches ChatGPT Plus, Pro, Business and Enterprise plans in the following days, along with the OpenAI API, Microsoft Azure and AWS Bedrock.

Is GPT-6 AGI?

There is no formal claim to that effect. OpenAI president Greg Brockman said it is “not unreasonable to feel that we are now in the AGI era”, which is an impression, not a measurement. The ARC Prize Foundation, which maintains the reasoning benchmark where the model performed best, writes explicitly that saturating that benchmark would not be proof of AGI, because its scope is narrow next to the real world.

Why does the same benchmark show 62.7% and 99.9%?

Because the harnesses differ. On ARC Prize’s standard harness, identical for every model, Astra scored 62.7%. With a provider-specific adapter that preserves reasoning state between calls and lets the model reuse earlier work, it scored 99.9%. Both numbers are real and measure different things: the second one includes the scaffolding, not just the model.

How much does GPT-6 Astra cost?

In the API, 10 dollars per million input tokens and 50 dollars per million output tokens, with cached input at 1 dollar. Fast mode costs twice as much and Batch and Flex modes half. That is roughly 2.5 times the price of the previous generation, GPT-5.6 Sol.

Should you switch models now?

Only after measuring. Run the tasks your team repeats most on both models and compare cost per completed task and rework rate, not price per token. A model that costs more per token can be cheaper per task if it finishes in fewer steps, and the reverse happens too.

What changes for my company’s security?

It is the first OpenAI model rated Critical for cyber capability under the company’s own framework, and it discovered two previously unknown vulnerabilities during an evaluation. The public version refuses offensive tasks and accepts defensive ones. In practice, finding flaws became cheaper for both sides, and the bottleneck becomes your patching and update routine.