OpenAI has called GPT-6 Astra its most intelligent model yet and the one most closely aligned with human intentions. It achieved near-perfect results in challenging mathematics, reasoning, and cybersecurity tests, but independent experts are not yet ready to recognize it as full-fledged artificial general intelligence.
The new model is the result of years of OpenAI research into pretraining, reinforcement learning, and AI alignment. The company also used substantial computing resources, although it has not disclosed their exact scale.
According to OpenAI, Astra scored 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench. These results have fueled claims that the new system could represent a major step toward artificial general intelligence, or AGI.
Nvidia CEO and OpenAI investor Jensen Huang went even further, declaring that Astra had already achieved AGI. This refers to a level at which a system can match or surpass human capabilities across a broad range of intellectual tasks.

At the same time, there is no definitive way to verify such a claim. Experts emphasize that there is still no universally accepted definition of AGI, while strong performance on individual benchmarks does not prove the existence of general intelligence.
Further questions emerged following a post by OpenAI Chief Scientist Jakub Pachocki. Several days after Astra’s launch, he warned that AI labs had created a form of intelligence whose inner workings they could no longer fully understand. As these systems continue to scale, the problem is likely to become even more pronounced.
The experts interviewed do not dispute Astra’s technical progress, but they believe the bold claims surrounding the model also serve a marketing purpose. Quickbase Chief Marketing Officer Alys Reynders noted that praise from the head of the world’s most valuable public company and one of the biggest beneficiaries of the AI boom carries considerable value for OpenAI. In her view, the current hype resembles a campaign designed to sustain interest among investors and other participants in the AI race.
Russell Twilligear, Head of AI Research and Development at BlogBuster, also does not believe that full-fledged AGI has already been achieved. However, he is convinced that OpenAI has come closer to reaching that level than other major AI labs, including Anthropic. “At this point, they are literally knocking on the door,” Twilligear said.
Experts also point to Nvidia’s financial interest. If artificial general intelligence is genuinely achieved, demand for the company’s accelerators and other infrastructure could rise significantly.
FrontierMath Tier 4 evaluates a model’s ability to solve the most difficult research-level mathematics problems. Astra scored 98%, while GPT-5 achieved approximately 83% at maximum compute.
ARC-AGI-3 assesses how well a system handles unfamiliar tasks that were not included in its training data. Astra’s score of 99.9% means the benchmark has effectively reached saturation: it can barely distinguish between Astra and a hypothetically more capable system.
For comparison, a human scored 48% on this benchmark, while GPT-5.6 Sol is estimated to score around 30%.
ARC Prize Foundation President Greg Kamradt said Astra surpassed the human action-efficiency baseline on 96% of all levels and effectively achieved human parity on the benchmark. He described it as the best model the organization had ever tested and a significant leap in frontier-model performance.
However, even such a result applies only to a specific testing environment. It does not prove that the model can perform every possible real-world task at a human level.
ExploitBench evaluates whether a model can trigger software bugs, build exploit primitives, reach vulnerable code, and achieve arbitrary code execution.
Astra was initially tested without the safeguards used in its public version. Researchers evaluated it on ExploitBench and ExploitGym to determine whether the system could turn known vulnerabilities into working exploits.
Because of concerns that benchmark data might have entered the model’s training set, OpenAI created an internal version of ExploitBench. Astra scored 100% on it, compared with 78.5% for GPT-5.6 Sol. On ExploitGym, the new model achieved 42.4%, while Sol scored 30.3%.
Between June and August 2026, ExploitBench included 20 high-severity vulnerabilities in the V8 engine across 13 stable Chrome releases. The model’s task was to achieve arbitrary code execution in V8 and official Linux versions of Chrome.
OpenAI notes that, under the evaluation’s constraints, some vulnerabilities may not permit arbitrary code execution at all. This means that a perfect result could theoretically be unattainable, making Astra’s reported 100% score particularly notable.
Technology companies regularly use benchmark results to demonstrate that their new models outperform competitors. However, the saturation of individual benchmarks creates another problem: once a system scores close to 100%, the test is no longer useful for comparing it with even more capable models. OpenAI claims that Astra not only achieved near-perfect results but also helped solve open problems in mathematics. Nevertheless, these claims still require independent verification, while researchers will need to develop more demanding benchmarks.