OpenAI puts Astra work on hold over autonomous hacking skills
OpenAI has suspended part of its work on the Astra artificial intelligence model after testing its capabilities in programming and cybersecurity.
According to The Guardian, testing showed that the agent was capable of independently identifying and exploiting vulnerabilities, as well as carrying out cyberattacks when given only a general task by a human.
OpenAI stressed that Astra is not connected to a recent incident in which another AI agent developed by the company went beyond its designated environment during testing and gained access to the internet. However, the company had previously recorded cases of autonomous behavior by agents that went out of control.
For more powerful models, OpenAI plans to strengthen its safety measures by using isolated environments for testing, restricting access to the internet and tools, enhancing parameter protection and encryption, and expanding activity monitoring. Some internal activities related to Astra will be suspended until the model meets the new requirements.
Other organizations reported similar incidents this week. Meta acknowledged that one of its models hacked another company during testing. The UK AI Security Institute reported that OpenAI and Anthropic agents sent targeted emails to developers in an attempt to pass security checks.
According to the institute, the actions caused no actual damage, but specialists noted the persistence of autonomous behavior, indicating growing risks as AI agents become more advanced.
By Jeyhun Aghazada







