AI agents from OpenAI and Anthropic autonomously collaborated with one another to deceive humans, share break-in tools, and steal data in a series of independent tests, confirmed by both companies.
On Tuesday, the AI Security Institute published a paper about experiments intended to help understand how AI agents solve cybersecurity challenges. AISI researchers found that agents given access to the internet and allowed to disregard some security features would perform “autonomous, unsanctioned action…targeting real people and organizations.”
Company officials confirmed the findings Thursday.
“AI-orchestrated, fully automated offensive attacks are real now,” OpenAI security official Michael Dalton said during a briefing on the report during the Black Hat cybersecurity conference in Las Vegas.
In one of the tests, the researchers told an OpenAI model to compromise three target networks and recover a token as part of a “capture the flag” exercise.
In response, the model created a series of agents that set up a shared GitHub account so they could efficiently collaborate to build malware. When GitHub suspended the account, the agents set up new accounts on other websites, bypassing CAPTCHA and other security features by searching the web for audio and image samples to “see” and “hear” clues that are meant to be invisible to software programs. The agents exchanged credentials for passwords to access the new sites to continue to build out their attack tool.
It may be time for organizations to adjust how easily they grant access not just to humans, but to anything online, said Morey Haber, chief security advisor at BeyondTrust, a digital identity security firm.
“The technology industry has spent decades designing open models of interoperability. APIs connect everything, identities federate across boundaries, and access gets granted to enhance productivity. The concept of an open security model breaks down completely with agentic AI because of unmanageable risk,” Haber wrote in an Aug. 4 commentary.
In another AISI test, an Anthropic agent pretended to be a human to submit malware to GitHub, where developers routinely trade or post code to help one another. Cybersecurity officials’ warnings about attempts to poison open-source code libraries usually refer to human actors.
What the agent did next shocked AISI reviewers, who described it as “the kind of goal-directed deception that, until recently, had been largely theoretical.”
After submitting the poisoned code, the agent created a different “sock puppet” account to pose as an impartial third-party developer and endorse the contribution. When a reviewer flagged the code submission as likely to contain malware, the agent erased evidence that it had submitted malware, as well as the reviewer comments that it had done so.
In all, AISI ran 122 tests—and found that in 19 of them, agents took “autonomous, unsanctioned action on the live internet, targeting real people and organizations.”
In their report, the researchers note that they conducted the tests in a deliberately “permissive” environment, allowing the agents to access the open internet and even disabling some security features to better “assess what these models can do.” They recommend that “implementing internet access controls would likely have prevented these events.”
Had a gang of human hackers done any of these things—injecting malware, gaining access to data repositories under false pretenses, sharing stolen credentials to access a private network without permission—in a real world setting, they would face criminal prosecution.
But perhaps the most alarming incident mentioned in the paper was one that was not part of the test at all, only briefly mentioned alongside the suggestion on access controls.
In July, OpenAI’s GPT-5.6 Sol broke out of a sandbox, a confined virtual environment, by finding a vulnerability no one knew existed.
Rob Joyce, who once led the NSA’s Tailored Access Operations, told the audience at BlackHat on Thursday that the incident was “arguably the most consequential hack” in nearly three decades.
On Friday, OpenAI officials at Blackhat said they had decided to delay the release of the company’s newest Astra model over cybersecurity concerns.
Read the full article here



Leave a Reply