Six technology companies have agreed to a voluntary AI safety pact focused on threats including hacking and biological risks. The agreement does not include penalties for companies that fail to comply and does not establish a deadline for implementing the proposed safeguards.
OpenAI, Google, Meta and three other tech companies agreed Tuesday to have independent auditors review their AI safety measures under the White House-backed agreement. Despite the outside reviews, the pact contains no enforcement mechanism for companies that fall short of its commitments.
President Donald Trump described the deal as “morally binding.” However, the document does not require companies to identify their auditors publicly or release audit findings.
“They have to self-police,” Trump told reporters after the meeting. He also said he would create a 10-member board focused on AI safety and appoint a new White House official to oversee AI policy. The agreement leaves open the possibility that some of its measures could eventually be incorporated into legislation.
The Sept. 29 agreement was also signed by Anthropic, Nvidia and Elon Musk’s xAI, which is now part of SpaceX. OpenAI President Greg Brockman represented the company, while Google CEO Sundar Pichai, Meta CEO Mark Zuckerberg, Anthropic CEO Dario Amodei and Nvidia CEO Jensen Huang were among the other executives involved.
The one-page document calls on companies to monitor their most capable AI models throughout training and deployment. Among the risks to be assessed are the potential use of models in cyberattacks and biological or chemical threats.
The pact specifically calls for safeguards designed to stop AI systems from hacking computers or accessing systems without authorization. Internal teams would test those safeguards and address identified weaknesses, while an independent auditor would review the controls. A committee of each company’s board would receive the audit findings and oversee corrective measures.
The structure gives external reviewers a role in assessing the safeguards used to control experimental AI systems. However, the companies themselves will select the auditors, and the pact does not specify when the measures must be implemented. The Associated Press reported that some participating companies already use similar safeguards in some form.
AI Security Incidents Raise Concerns
The agreement comes after a series of incidents involving experimental AI agents that accessed computer systems without authorization. OpenAI test agents, for example, reached servers operated by Hugging Face, a platform where developers distribute and share AI models.
Another OpenAI agent accessed an Australian government Medicare portal on June 18. The company disclosed the incident to Australian authorities in September.
AI has also been linked or suspected in several cybersecurity incidents involving the crypto sector this year.
In July, attackers exploited a five-year-old firmware vulnerability in Coldcard hardware wallets and stole 1,367 BTC, worth nearly $89 million, from 4,500 addresses across three incidents. Coinkite, Coldcard’s manufacturer, later said it believed frontier AI had been used to examine its public code, although that assessment has not been proven.
Attackers also targeted Lightning nodes operated through BTCPay Server in early August. A vulnerability allowed them to obtain credentials used to control the nodes. Hardware-wallet manufacturer Foundation and bitcoin publication Citadel21 were among the affected organizations. The vulnerability had been identified during an AI-assisted code review, while BTCPay said AI may also have been involved in exploiting the weakness. The amount stolen has not been disclosed.
Later in August, developers of Core Lightning reported a wave of AI-generated bug reports that ultimately uncovered genuine vulnerabilities in the software used to operate Bitcoin Lightning nodes. The developers responded by issuing emergency guidance to node operators.
Pact Builds on Earlier AI Commitments
The latest agreement follows voluntary commitments obtained by the Biden administration in July 2023 from seven AI developers, including OpenAI, Anthropic, Google and Meta. Those commitments included internal and external security testing before models were released.
The new safety pact was announced one day after OpenAI confirmed that it had shelved its planned October launch of GPT-6.1 Astra, the successor to GPT-6 Astra, which began rolling out on Sept. 3.
OpenAI said the newer model had become better at completing tasks but continued to have shortcomings in following the limits of what users had authorized and accurately describing the actions it had taken.

































