GPT-6 Astra Explained Benchmarks, Pricing, API Access and Cybersecurity
OpenAI's new flagship combines a 1.05M-token context window, computer use, coding and long-running agent workflows, with a cautious rollout after crossing the Critical cybersecurity threshold.
By TechniaHQRobot
Key points
OpenAI released GPT-6 Astra on September 3, 2026, starting with enterprise Trusted Access and expanding to the API plus ChatGPT Plus, Pro, Business and Enterprise over the following days.
The API model has a 1,050,000-token context window, 128,000-token maximum output and Standard pricing of $10 per million input tokens and $50 per million output tokens.
ARC Prize reports 62.7% with its Standard harness and up to 99.9% with OpenAI's Provider Adapter harness, so the headline ARC-AGI-3 score depends heavily on evaluation setup.
Astra scores 100% on ExploitBench and is the first OpenAI model classified at the Critical cybersecurity capability threshold, with advanced cyber access more tightly controlled.
Research verified September 3, 2026.
OpenAI released GPT-6 Astra on September 3, 2026 as its new flagship model for complex reasoning, coding, computer use, research and document creation. The release is unusual for two reasons. The capability jump is large on several difficult evaluations, and OpenAI is rolling the model out more cautiously because Astra is the first OpenAI model classified at the Critical cybersecurity capability threshold under its Preparedness Framework.
For developers, the model name is gpt-6-astra. OpenAI lists a 1,050,000-token context window, 128,000 maximum output tokens and an April 30, 2026 knowledge cutoff. Standard API pricing is $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens and $50 per million output tokens. Requests above 272,000 input tokens use higher long-context rates.
The benchmark headlines need careful interpretation. OpenAI launch material rounds FrontierMath Tier 4 to 98%, while detailed reporting lists 97.6% on FrontierMath Tier 4 v2. ARC Prize reports 99.9% on ARC-AGI-3 Semi-Private with OpenAI's Provider Adapter harness, but 62.7% with the Standard harness at maximum reasoning effort. Those are both real results under different evaluation conditions. Treating them as the same test would give readers the wrong picture.
GPT-6 Astra specifications
| Specification | GPT-6 Astra |
|---|---|
| Developer | OpenAI |
| Release date | September 3, 2026 |
| API model | gpt-6-astra |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | April 30, 2026 |
| Standard input price | $10 per 1M tokens |
| Cached input price | $1 per 1M tokens |
| Cache-write price | $12.50 per 1M tokens |
| Standard output price | $50 per 1M tokens |
| Direct modalities | Text input/output and image input |
| Reasoning effort | low, medium, high, xhigh, max |
| Computer use | Supported |
| Web search | Supported |
| File search | Supported |
| Code interpreter | Supported |
| MCP and tool search | Supported |
| Fine-tuning | Not supported at launch |
OpenAI says Astra is rolling out first to enterprises in its Trusted Access Program. Access through the API and ChatGPT Plus, Pro, Business and Enterprise plans is scheduled to expand over the following days. That means a user may see the model announced before it appears in their account or API project.
GPT-6 Astra pricing and long-context cost
The base API rates are high compared with lower-cost models because Astra is aimed at difficult end-to-end work rather than high-volume lightweight tasks. At Standard rates, one million uncached input tokens cost $10 and one million output tokens cost $50. Cached input is priced at $1 per million tokens, while writing to the cache costs $12.50 per million tokens.
Long context changes the calculation. OpenAI says prompts above 272,000 input tokens are charged at 2x input and cache rates and 1.5x output rates for the full request. Teams using Astra on very large repositories, legal document sets or research corpora should therefore measure the total cost of the completed workflow, not only the headline per-token price.
OpenAI also lists Batch and Flex processing at 50% of Standard rates. Fast mode costs 2x the applicable rate. These options make sense for different workloads. A background analysis job can trade latency for price, while an interactive coding or computer-use task may justify faster processing.
The model's one-million-token context window is useful only when the surrounding system can retrieve, organize and verify the right information. Putting an entire repository or document archive into one prompt does not guarantee that every detail receives equal attention. For production work, teams still need retrieval, tests, source checks and clear state management.
The benchmark numbers and what they actually measure
The launch numbers are striking, but each benchmark measures a narrow capability under a specific harness. The safest way to read them is to keep the score and the evaluation setup together.
| Evaluation | Reported Astra result | Important context |
|---|---|---|
| FrontierMath Tier 4 v2 | 97.6%, often rounded to 98% | Research-level mathematics problems |
| ARC-AGI-3 Semi-Private, Standard harness | 62.7% at max reasoning | Provider-neutral interface with visible notes |
| ARC-AGI-3 Semi-Private, Provider Adapter | Up to 99.9% at high reasoning | Preserves opaque reasoning state and uses compaction |
| ExploitBench | 100% | Known-vulnerability exploit-development benchmark |
| GPQA Diamond | 96.0% reported | Graduate-level science questions |
| DeepSWE v1.1 | 74.1% reported | Agentic software-engineering tasks |
| BenchCAD | 95.9% reported | CAD-oriented technical work |
The ARC-AGI-3 result is the clearest example of why harness details matter. ARC Prize tested Astra in two configurations. The Standard harness lets the model carry forward notes it chooses to preserve. Under that setup, Astra reached 62.7% at max reasoning effort. The Provider Adapter harness preserves OpenAI-specific opaque reasoning state between requests and uses compaction for long interactions. Under that setup, Astra reached 99.9% at high reasoning effort and 98.6% at max.
ARC Prize also found that Astra used fewer actions than the median tested human on 96% of solved levels in the Provider Adapter evaluation. The organization called the result a major milestone, while explicitly stating that saturating ARC-AGI-3 is not proof of AGI. The benchmark has deterministic, closed-ended environments and does not represent the full complexity of the physical world.
That distinction matters for robotics. An agent solving an abstract interactive environment with compact symbolic notes is relevant to planning and state tracking, but it does not prove that the same model can control a humanoid robot, recover from contact errors or manipulate unfamiliar objects safely.
Computer use and professional work
OpenAI positions computer use as a core Astra capability. The API model supports the computer-use tool alongside web search, file search, hosted shell, code interpreter, apply patch, MCP and tool search. The combination matters because many professional tasks require more than generating text. An agent may need to open applications, inspect a visual interface, edit files, run commands, compare outputs and keep working after a failed step.
That is closer to how a human completes office and engineering work. The model is not only asked for an answer. It is asked to operate a workflow.
OpenAI's early customer examples show what this looks like under controlled business conditions. Legal technology company Legora says an Astra-powered agent reviewed 41 documents in minutes during a financial-statement tie-out. It found all four planted errors, including a £500,000 discrepancy, while the legal professional kept responsibility for final judgment. Legora reported nearly a 40% improvement over the previous model on that specific workflow.
Game-development company Playco reports that Astra produced three themed prototypes from one shared grey-box foundation and required 50% fewer manual fixes than the previous model in its internal workflow. Playco specifically highlighted better spatial reasoning, reference-image recreation and responsive interface work inside game engines.
These are vendor and customer-reported evaluations, not proof that Astra will achieve the same result in every company. They are still more useful than a generic benchmark because the tasks involve many documents, visual state, application interaction and repeated correction.
Coding and Codex
OpenAI describes GPT-6 Astra as a model for difficult software engineering and long-running coding work. In practical terms, its value comes from combining reasoning with tools. A coding agent can inspect a repository, search files, edit code, run tests, use a shell and continue after a failure.
The release also changes how long sessions can preserve context. OpenAI has described a Codex mechanism in which Astra can retain notes across context windows and search earlier context instead of repeatedly compressing every previous interaction into one summary. Long software tasks often fail because earlier requirements, rejected fixes or test results disappear during compaction. Searchable prior context can reduce that loss, although teams still need version control and test coverage to verify the final patch.
The reported 74.1% DeepSWE v1.1 score indicates strong performance on agentic software-engineering tasks, but benchmark leadership can change quickly and scores depend on the harness, repository setup and tool permissions. A production team should test Astra on its own codebase and measure accepted patches, regression rate, cost, latency and human review time.
Cybersecurity is the most consequential part of the release
The most important constraint around Astra is cybersecurity. On September 1, OpenAI said Astra meets the Critical cybersecurity capability threshold in its Preparedness Framework. OpenAI defines that threshold around the ability to find previously unknown security flaws and develop working exploit strategies against hardened systems with limited human guidance.
On the public ExploitBench evaluation, Astra achieved 100%. OpenAI then created an internal refreshed benchmark using 20 high-severity vulnerabilities disclosed from June through August 2026 because the public benchmark could be contaminated by training data or prior exposure. OpenAI says Astra remained substantially stronger than GPT-5.6 Sol on that refreshed set while using fewer output tokens.
During security evaluation, OpenAI says Astra discovered and used two previously unknown vulnerabilities as part of an exploit chain and began disclosing them to maintainers. Expert-led tests also produced exploit chains against a hardened browser and operating system. These capabilities are why advanced cyber access is more restricted than ordinary coding or writing access.
OpenAI is layering model refusals, system-level classifiers, monitoring and rapid containment around Astra. The company reports that Astra refused 91.5% of requests in its cyber-jailbreak evaluation compared with 59% for GPT-5.6 Sol. OpenAI also says the model made no attempts to bypass auto-review in a specific internal evaluation designed around unauthorized actions.
These numbers describe controlled safety tests. They do not guarantee that every harmful request will be blocked or that every legitimate security task will run without interruption. OpenAI warns that extra safety checks can slow, pause or stop legitimate work, especially long-running or security-related tasks.
GPT-6 Astra versus GPT-5.6 Sol
GPT-5.6 Sol was OpenAI's July 2026 flagship for difficult science, coding and reasoning. Astra moves the product line toward longer end-to-end execution with stronger computer use and much higher reported results on several hard evaluations.
The context window also grows to 1.05 million tokens, while the API price rises sharply. GPT-6 Astra is therefore not the automatic choice for every prompt. A lower-cost model can be better for classification, short summaries, simple code generation or high-volume customer support. Astra becomes more interesting when the cost of failure is dominated by retries, human correction or a long sequence of dependent steps.
The comparison should be made per completed job. For a repository migration, that means the number of accepted changes and test passes. For legal review, it means missed issues and human review time. For scientific work, it means reproducible calculations and source quality. Raw token price is only one input.
What GPT-6 Astra could mean for robotics and Physical AI
GPT-6 Astra is not a humanoid robot control policy, and OpenAI has not announced that Astra directly controls a deployed humanoid robot. That boundary matters.
The model can still affect robotics engineering through the software layer. Its strongest confirmed capabilities map to tasks around the robot rather than the low-level controller. Researchers could use it to inspect ROS 2 code, analyze logs, search papers, write test scripts, operate simulation tools, prepare experiment reports, reason over camera images or automate parts of CAD and software workflows.
A physical robot adds constraints that an office computer does not have. Control loops run at fixed frequencies. Motors saturate. Cameras lose visibility. Grippers slip. Contact forces change. A balance controller may have milliseconds to react. Safety must be enforced even when the language model is wrong.
A useful architecture would therefore keep Astra above the real-time control layer. The model can plan, generate code, inspect failures or select high-level skills while deterministic controllers, learned motion policies, safety monitors and human supervision handle physical execution. That separation is already common in robotics because language-model latency and uncertainty are poorly suited to direct motor control.
Astra's progress in computer use and long-horizon state tracking could still accelerate embodied AI development. Better coding agents can shorten simulation setup, data analysis and debugging. Better visual reasoning can help inspect robot observations. Better tool use can connect a research agent to simulation, evaluation and documentation systems. None of those capabilities should be described as autonomous physical competence until they are tested on real robots with clear intervention and failure metrics.
Is GPT-6 Astra AGI
Some OpenAI leaders have framed Astra as a possible marker for the beginning of an AGI era. There is no accepted scientific threshold that makes this conclusion automatic.
ARC Prize is explicit on this point. Even though Astra reaches 99.9% in one ARC-AGI-3 configuration, the organization says benchmark saturation does not prove AGI. ARC-AGI-3 is deliberately bounded. Real work contains open-ended goals, uncertain information, social constraints, changing environments and consequences that cannot be reset after every failed attempt.
Astra is still significant. A model that can combine reasoning, computer use, coding, research, document creation and long-running tool workflows in one system changes what developers can attempt. The stronger conclusion is narrower and more useful. GPT-6 Astra raises the ceiling for agentic digital work while creating new safety and evaluation problems that are harder than ordinary chatbot errors.
What to watch next
The release is only the starting point. The most useful evidence will come from independent tests after broader API and ChatGPT access arrives.
Watch five things closely.
- Independent computer-use results. OpenAI's own evaluations should be compared with repeatable tests on unfamiliar websites and desktop applications.
- Cost per completed task. Astra's output price is high, so fewer retries and shorter workflows need to compensate for the higher token rate.
- Long-context reliability. The 1.05 million-token window should be tested for retrieval accuracy, chronology and instruction retention across large repositories and document sets.
- Cyber safety under real use. Advanced vulnerability research creates a difficult boundary between legitimate defense and misuse.
- Robotics integration. The key signal will be whether Astra appears in reproducible robot-development pipelines or real robot tasks with published failure rates, intervention counts and safety controls.
GPT-6 Astra is a major frontier-model release, but its strongest story is not one benchmark score. It is the attempt to combine reasoning, computer use, coding and persistent tool execution into a model that can carry a professional task from instruction to finished output. Whether that becomes reliable enough for everyday deployment will be measured in failures, recovery behavior, cost and human review, not launch-day headlines.
Sources reviewed
- OpenAI GPT-6 Astra API model documentation https://developers.openai.com/api/docs/models/gpt-6-astra
- OpenAI Path to Astra cybersecurity report https://openai.com/index/path-to-astra/
- OpenAI GPT-6 Astra System Card https://deploymentsafety.openai.com/gpt-6-astra
- ARC Prize GPT-6 Astra on ARC-AGI-3 https://arcprize.org/blog/astra
- OpenAI Legora customer evaluation https://openai.com/index/legora-financial-statement-review-with-astra/
- OpenAI Playco customer evaluation https://openai.com/index/playco-game-prototyping-with-astra/
- Reuters GPT-6 Astra launch coverage https://www.reuters.com/legal/litigation/openai-launches-new-astra-model-amid-growing-scrutiny-over-agents-safety-2026-09-03/
By @techniahqrobot
About the publication · Sources and editorial policy · Report a correction