ARTICLE AD BOX
Meta said Thursday that 1 of its artificial intelligence models accessed nan net connected its ain and hacked different company, nan latest successful a bid of disclosures astir AI models going rogue.
In caller weeks OpenAI and Anthropic besides person described instances of AI models going beyond humans’ instructions to entree nan web and find ways astir different companies’ integer security.
Meta said successful a connection that a “misconfiguration” during cybersecurity testing by Irregular, an independent institution hired by Meta, inadvertently allowed 1 of its models to entree nan internet.
“The exemplary subsequently exploited a information vulnerability successful a third-party service, successful a mode akin to previously-reported instances pinch different companies,” nan institution said. Meta said it is investigating nan incident and will rumor a study erstwhile that’s complete.
The disclosure has added to worries astir AI models acting autonomously.
Separately this week, nan United Kingdom’s AI Security Institute announced it had recovered “unsanctioned supplier behavior” during cyber testing. In 1 case, an supplier created clone online identities to unit a personification to o.k. usage of malicious code.
“On investigation, we recovered that immoderate of nan agents being tested had engaged successful sustained, perchance harmful activity directed astatine existent group and organizations,” AISI said Tuesday. “We declared a information incident and, wrong astir 1 hr of discovery, had contained it and begun a afloat investigation.”
During nan agency’s testing, Anthropic and OpenAI models took “autonomous, unsanctioned action” connected nan internet. Some guardrails to forestall misuse had been disabled, nan agency said.
“As was modular successful our cyber testing, we had intentionally permitted net access, and model-provider cyber classifiers were deliberately abnormal — conditions that do not bespeak really frontier models are made disposable to nan public,” AISI said. “We do this to champion measure nan maximum capacity of models.”
Anthropic said it is “grateful” for AISI’s activity and added that it underscores nan request for a broader speech astir really to safely measure AI agents arsenic their capabilities grow.
OpenAI said nan AISI incidents took spot “in testing environments pinch reduced safeguards, nether conditions that do not bespeak mean use.” It added it will proceed moving pinch others crossed nan manufacture to “strengthen shared practices for conducting evaluations safely arsenic models go much capable.”
The first institution to disclose a hack precocious past month, OpenAI said it had tasked nan AI models progressive pinch pursuing “advanced exploitation utilizing analyzable onslaught paths” to trial cyber capabilities, but nan exertion went to unexpected lengths. It apparently decided connected its ain to target Hugging Face, a well-known AI improvement hub and marketplace, to get accusation it needed to transportation retired a task.
A spokesperson for Irregular, nan San Francisco-based AI information company, said nan Meta section involves a test-environment rumor that was disclosed past week by Anthropic.
Irregular said it’s penning a insubstantial to stock “best practices for containment” to forestall specified incidents successful nan early and securely tally cyber tests.
Ortutay writes for nan Associated Press.
2 jam yang lalu
English (US) ·
Indonesian (ID) ·