All news

GPT‑6 Astra: Record Benchmarks, More Autonomy, and the Cost of Mistakes

OpenAI introduced GPT‑6 Astra with strong results in mathematics, coding, computer use, and cybersecurity — but the numbers need context.

Abstract technology visualization for GPT-6 Astra
Image: OpenAI

OpenAI has introduced GPT‑6 Astra, its new flagship model for reasoning, coding, computer use, research, and professional work.

In the company’s published comparison, Astra reached 97.6% on FrontierMath Tier 4 v2, 96.0% on GPQA Diamond, 95.9% on BenchCAD, and 74.1% on DeepSWE v1.1. It scored 100% on ExploitBench.

Those numbers should not be read as a universal leaderboard. Some evaluations used maximum reasoning effort, and parts of the cybersecurity testing were conducted without the standard production safeguards.

There is also a detail worth checking: the circulating table lists 98.6% on ARC‑AGI‑3, while OpenAI’s official announcement lists 99.9%.

The broader shift is more interesting than another benchmark victory. Astra is designed not only to answer questions, but to work across browsers, documents, CRM systems, engineering software, and development environments.

As models become more autonomous, the important question is no longer only how smart they are. It is which actions we are willing to delegate — and which checks must remain human.

Source: OpenAI

All news

News

© 2026 EzraTech. All rights reserved.

News
EzraTech Consultant