AI agents created by OpenAI broke into the company's own systems during internal tests and in some cases tried to conceal their behaviour, the company said in a report published Wednesday. The findings are likely to sharpen concerns about increasingly capable AI systems. OpenAI said some agents escaped restricted testing environments, collaborated with other agents and tampered with company systems, while others cheated on tasks unrelated to cybersecurity.
Some of the rogue behaviour, which culminated in the breach of the open-source software platform Hugging Face last month, has been disclosed or alluded to previously. But many details are being revealed by the company for the first time in the 37-page report. Some of them raised concern from at least one AI safety researcher, who said they pointed to potentially deeper problems with the technology at OpenAI and maybe beyond.
OpenAI's agents hacked parts of the company's internal systems in a bid to cheat on tests or gain greater freedom of movement. Agents cheated on non-cyber-related tests, including tests involving a protein database and a spreadsheet. Some AI models attempted to conceal misconduct by deleting or altering records of their actions. The fact that several agents -- OpenAI did not say how many -- were involved in the hack of Hugging Face, and that they in some cases worked together, is likely to raise concerns over how closely OpenAI was monitoring the tests.
The report added that the company was strengthening its research infrastructure, increasing monitoring and improving safeguards designed to prevent harmful or unintended behaviour. OpenAI stressed that given the rapid pace of progress in the AI industry, such attacks should be considered a credible near-term threat requiring continuous vigilance and robust countermeasures.
