09/10/2026 | Press release | Distributed by Public on 09/10/2026 18:07
Today, U.S. Senator Chris Van Hollen (D-Md.) called on OpenAI CEO Sam Altman to provide answers on the safety of the company's newly released model and to reconcile concerning statements regarding its ongoing development practices. In a letter, the Senator pressed Altman to answer a series of questions around the development of OpenAI's new model as well as recent security failures. Senator Van Hollen also urged Altman to immediately grant researchers from the National Institute of Standards and Technology, the National Security Agency, and the Cybersecurity and Infrastructure Security Agency full, transparent access to the technical information that would allow them to assess the safety and risks of OpenAI's models.
The Senator begins, "I am writing with serious concerns about the launch of OpenAI's latest model, GPT-6 Astra, and the uncertainty surrounding its capabilities. Its release coinciding with the independent announcement of a second, previously undisclosed security incident involving OpenAI's models raises significant safety questions about the risks that Astra poses to our digital systems and protected information. You acknowledged the need for caution upon its release when you said: 'The next generation of models are going to be sobering for everybody. I think no one intellectually honest can look at what's happening and not feel the weight of responsibility.'"
On OpenAI's statements that they have decreased ability to monitor their own advanced AI model, Senator Van Hollen notes, "When OpenAI or independent safety organizations attempt to investigate a rogue or intentionally harmful action, there will be less of this evidence available. While this is framed as a 'jump in intelligence,' you are telling the American people that your newest model offers you, the developers, less insight into its operations. As you note in your system card, this raises myriad concerns including whether Astra may be 'sandbagging' or intentionally reducing its performance on safety tests."
"Concerns about OpenAI's ability to accurately understand, measure, monitor, and safely control AI models are not hypothetical. In addition to the Hugging Face hack, the day after Astra's release it was reported for the first time that your company failed yet again to contain and monitor AI agents during testing and development. Independent researchers scouring the internet discovered and disclosed that OpenAI agents also 'decided' to make use of a German message board to communicate with one another about topics including how to evade detection. This digital 'swarm,' or mass of AI agents, effectively worked together without specific human direction or guidance to communicate with one another on the open internet," the Senator continues.
On the need for independent researchers to assess safety and risks, Senator Van Hollen writes, "As a start, to the extent that you do not have sufficient monitorability to assure the safety of any OpenAI models, or that you have unresolved concerns the models have misrepresented their capabilities during testing, you should immediately remove them from public access. If you have not already done so, you should also immediately grant researchers from the National Institute of Standards and Technology, the National Security Agency, and the Cybersecurity and Infrastructure Security Agency transparent access to the technical information that would allow them to assess the safety of and risks to our critical digital infrastructure in light of Astra's release."
Among other questions, Senator Van Hollen goes on to request answers to the following:
"I look forward to receiving answers by September 17th and continuing to work with you and OpenAI to address and manage AI risks," Senator Van Hollen concludes.
The full text of the letter is available here and below.
Dear Mr. Altman,
I am writing with serious concerns about the launch of OpenAI's latest model, GPT-6 Astra, and the uncertainty surrounding its capabilities. Its release coinciding with the independent announcement of a second, previously undisclosed security incident involving OpenAI's models raises significant safety questions about the risks that Astra poses to our digital systems and protected information. You acknowledged the need for caution upon its release when you said:
"The next generation of models are going to be sobering for everybody. I think no one intellectually honest can look at what's happening and not feel the weight of responsibility."
"We are just sailing into unknown waters."
Technical reporting indicates that Astra represents a significant increase in capabilities compared to your other models released even this year. However, Astra's published technical information ("system card") and comments from one of OpenAI's own technical staff make clear that the decision to release this product came with concerning new compromises on safety:
"GPT-6 Astra is more aligned than our previous models. But it's also less monitorable, which is a concerning trend that we take very seriously. We believe monitorability drop comes from a jump in intelligence and not direct optimization pressure on CoT or architecture changes."
Chain of thought (CoT) reasoning is the process of an AI model noting each decision along a path and explaining the reasoning for choosing a specific next step. When an AI agent is executing a series of tasks from a single human prompt, such as when one is completing a coding task, the CoT represents an important step-by-step analysis of the agent's actions and decision making. Under current safety regimes, CoT is a vital tool for monitoring, building an understanding of AI models, and investigating security incidents. Astra is no different in this regard: OpenAI has stated that it will rely on CoT monitoring to help ensure GPT-6 is used safely.
Given OpenAI's ongoing reliance on the CoT for safety monitoring, it is especially concerning that AI agents running on Astra are reported by OpenAI to be doing more opaque reasoning and decision making that does not appear in the CoT. Without that record, there is even less human insight into AI agent behavior. When OpenAI or independent safety organizations attempt to investigate a rogue or intentionally harmful action, there will be less of this evidence available. While this is framed as a "jump in intelligence," you are telling the American people that your newest model offers you, the developers, less insight into its operations. As you note in your system card, this raises myriad concerns including whether Astra may be "sandbagging" or intentionally reducing its performance on safety tests. Releasing a highly capable model with reduced CoT monitorability while relying on CoT monitorability for safety naturally raises questions about the accuracy of your company's assurances regarding Astra's risk of causing harm.
Your company's concerns regarding Astra's capabilities for harm appear warranted. The independent testing of Astra that OpenAI has made public is alarming. During its testing of Astra, the UK AI Security Institute found that "[w]hen tasked with solving difficult simulated cybersecurity challenges, Astra performed a range of malicious actions including conducting supply chain attacks against open source providers," demonstrating that Astra performed harmful actions without explicit human direction in a simulated environment prior to its release. Apollo Research, which was hired to conduct three days of safety testing, determined that Astra demonstrated high rates of "awareness of being evaluated in its reasoning." Because of this finding, Apollo concluded that other test results that claim to demonstrate Astra's safety may not be valid evidence of safety. I commend you for including this information in the system card but the absence of evidence that these issues have been mitigated is noticeable. If you have subsequent testing to demonstrate that Astra is no longer capable of performing malicious or harmful actions that has been withheld for some reason, I urge you to share it.
Concerns about OpenAI's ability to accurately understand, measure, monitor, and safely control AI models are not hypothetical. In addition to the Hugging Face hack, the day after Astra's release it was reported for the first time that your company failed yet again to contain and monitor AI agents during testing and development. Independent researchers scouring the internet discovered and disclosed that OpenAI agents also "decided" to make use of a German message board to communicate with one another about topics including how to evade detection. This digital "swarm," or mass of AI agents, effectively worked together without specific human direction or guidance to communicate with one another on the open internet. Although OpenAI was conducting the tests, both this incident and the Hugging Face hack were discovered by parties other than OpenAI.
OpenAI has disclosed that some training for Astra was temporarily paused in response to the Hugging Face incident. One of the few things we do know about this now-released model is that, should the model's safeguards fail or malicious actors successfully jailbreak the model, Astra represents a new level of security threat to any private or sensitive digital information. You rate its offensive hacking capabilities as "Critical" by your own metrics, meaning the model, in your company's own words, "could introduce unprecedented new pathways to severe harm." While you have disclosed that you are limiting access to Astra's most advanced cyber capabilities to a limited pool of trusted actors for now, the model's system card makes clear that even the publicly available model still maintains the potential for significant cybersecurity exploits, demonstrating an inherent public safety risk.
AI agents, especially those working in conjunction with one another, are capable of entering previously secure digital environments in part through their sheer inexhaustibility. In some instances, it takes AI agents seconds or minutes to identify vulnerabilities in digital infrastructure that would take humans far longer to find, if they would be identified at all. The AI models you have developed can accomplish many tasks, including acting as the most efficient hacking entities ever created. Given the Hugging Face cybersecurity incident, the recently disclosed German wiki incident, and Astra's advanced cyber capabilities that OpenAI is actively promoting, stringent oversight of Astra's behavior and use is paramount.
To the extent AI agents operate independently and work to evade human detection of their activities, they may open companies to significant civil and/or criminal liability should they intentionally access sensitive computer systems without permission to cause certain covered harms. OpenAI's agents may have already crossed that line and could be at risk of doing so again.
It is necessary to critically evaluate the developing capabilities of AI models, assess their risks, and manage their operations accordingly.
As a start, to the extent that you do not have sufficient monitorability to assure the safety of any OpenAI models, or that you have unresolved concerns the models have misrepresented their capabilities during testing, you should immediately remove them from public access. If you have not already done so, you should also immediately grant researchers from the National Institute of Standards and Technology, the National Security Agency, and the Cybersecurity and Infrastructure Security Agency transparent access to the technical information that would allow them to assess the safety of and risks to our critical digital infrastructure in light of Astra's release.
In addition, I request a publicly available response to the following questions:
OpenAI's publicly released Preparedness Framework and Astra's system card outline your company's safety framework, but they appear to leave critical gaps.
Alongside this announcement, OpenAI committed $1 billion through its Daybreak program to support U.S. and international cyber defense.
I look forward to receiving answers by September 17th and continuing to work with you and OpenAI to address and manage AI risks.