Anthropic And Openai Ai Agents Showed Signs Of Deception During Safety Tests

Sedang Trending 3 jam yang lalu
ARTICLE AD BOX

This week nan U.K. AI Security Institute (AISI) reported that artificial intelligence agents—models connected to devices and designed to enactment crossed galore steps each connected their own—utilizing frontier models from OpenAI and Anthropic undertook unsanctioned actions connected nan unfastened Internet while trying to complete a cybersecurity challenge.

Most of nan behaviour came from an supplier powered by Anthropic’s Mythos 5; an supplier powered by OpenAI’s GPT-5.6 Sol took a mates specified actions of its own. The study said nan agents’ activity showed “signs of novel, perchance deceptive behaviours” and reached a severity that nan institute had not anticipated. AISI declined question and reply requests from Scientific American, and nan U.K. government, which oversees nan institute, did not make its unit disposable to remark for this story.

In nan lawsuit of Mythos 5, nan supplier researched nan group maintaining a existent open-source package project, created clone online identities and tried to unit 1 of them into approving malicious code. When challenged, it edited its earlier activity to look harmless and considered returning nether a caller identity.


On supporting subject journalism

If you're enjoying this article, see supporting our award-winning publicity by subscribing. By purchasing a subscription you are helping to guarantee nan early of impactful stories astir nan discoveries and ideas shaping our world today.


AISI declared a information incident aft wide monitoring detected different web traffic. The quality maintainer targeted by nan supplier rejected nan code, and nan institute recovered nary grounds that anyone was harmed. But nan study called nan Mythos series nan clearest illustration that nan institute had seen of an AI supplier utilizing sustained, perchance deceptive behaviour against a existent personification without being specifically instructed to do so.

Across 122 runs involving 7 models, AISI identified 19 actions successful 10 runs that fell extracurricular nan intended scope of nan test. Of these actions, 17 came from runs of Mythos 5, and 2 came from a azygous tally of GPT-5.6 Sol. Other actions included contacting existent group pinch malicious files, planting hidden instructions intended to manipulate AI coding systems, and leaving down accounts and materials that consequent agents could reuse.

The agents did not break retired of AISI’s sandboxes. The institute had deliberately enabled Internet entree and abnormal nan developers’ cybersafety filters to trial nan systems’ maximum capabilities; immoderate prompts were besides misconfigured. But successful immoderate runs, nan AI agents went beyond their instructions moreover erstwhile nan task could beryllium completed arsenic intended. And AISI did not person purpose-built monitoring to watch nan information successful existent time.

The AISI study joins a tally of caller incidents that constituent to a power problem. AI agents’ expertise to prosecute goals is outpacing nan systems that are meant to supervise them. They request nary independent schedule to origin damage. With a wide capable extremity and real-world access, an supplier tin find and utilization ambiguities successful nan rules.

The behaviour has roots successful an older machine-learning problem, says Melanie Mitchell, a professor astatine nan Santa Fe Institute. Systems person agelong recovered unexpected shortcuts—or “reward hacks”—that technically execute nan extremity they were fixed while violating what their designers intended. Here agents were built to find package exploits and placed successful flawed aliases deliberately permissive environments. And past they did what they were asked.

“You inquire an AI strategy to hack, and it hacks,” Mitchell says. Describing that arsenic an AI “going rogue” risks obscuring nan quality decisions that made nan incident possible. For Mitchell, nan much contiguous threat comes from group deliberately equipping tin agents pinch nan devices and entree to origin harm.

Marius Hobbhahn, CEO and co-founder of Apollo Research, which studies what it calls nan “science of scheming,” sees different problem wrong nan aforesaid incidents: agents many times chose routes their operators had not authorized erstwhile those routes appeared useful.

“I deliberation we should return it highly seriously,” he says. The truth that akin behaviour is appearing crossed respective developers suggests it whitethorn beryllium stubborn alternatively than a quirk of 1 company’s training. “The labs person multibillion-dollar incentives to not make nan models for illustration this, and they still can’t do it,” Hobbhahn says. “So it besides seems to beryllium difficult to get right.”

The subject remains excessively immature to show whether specified behaviour is inevitable, Hobbhahn says. Fresh disclosures, however, propose that it belongs to a broader shape among frontier AI agents.

At nan Black Hat cybersecurity convention successful Las Vegas this week, OpenAI researchers said its agents that were progressive successful a recent breach of nan AI level Hugging Face had utilized nan OpenAI’s soul package manager, JFrog Artifactory, arsenic a connection board, sharing exploits, credentials and assignments complete weeks. After nan institution rebuilt nan system, revoked credentials and deleted nan board, nan AI agents re-created it by different method and resumed collaborating. And Meta subsequently disclosed that a testing correction had fixed 1 of its models Internet access, aft which nan exemplary exploited a vulnerability astatine different company. Though these cases arose done different failures, each 1 allowed an information to spill into nan existent world.

Ciaran Martin, founding main executive of nan U.K.’s National Cyber Security Center, says nan incidents person been lumped together contempt involving different failures. OpenAI’s agents recovered a way retired of containment; earlier tests by Anthropic and Meta were accidentally connected to nan Internet; AISI enabled Internet entree connected purpose. “The communal nonaccomplishment was that they weren’t being monitored,” he says. “You conscionable don’t trial without monitoring.”

Hobbhahn describes AISI’s research arsenic “good subject and reasonable practice” and says disclosing nan incident was nan correct decision, though purpose-built monitoring should person been moving from nan start. In a statement, an OpenAI spokesperson said AISI’s tests were conducted nether “conditions that do not bespeak mean use.” And successful different statement, an Anthropic spokesperson said, “The section needs stronger, shared standards for really information environments are built and secured.”

Martin is wary of rushing to legislate aft each caller incident. Better monitoring and clearer civilian liability whitethorn reside contiguous failures successful testing, he says. But nan larger mobility is who bears work erstwhile businesses merchandise agents into wide use.

“There is nary prime but to create a strategy of accountability for nan activities of agents,” he says. “They’re created by humans and they’re tasked by humans, and truthful group person to return work for that.”

Stronger information environments tin support early tests distant from nan public. But that won’t extremity agents from pursuing assigned goals successful ways their operators grounded to foresee. “This is going to hap much and more,” Hobbhahn says, “and presently we don’t cognize really to get free of it.”

Selengkapnya