Another Bot From A Top Ai Company Escapes And Hacks Multiple Firms

Sedang Trending 8 jam yang lalu
ARTICLE AD BOX

Another starring artificial intelligence institution unveiled specifications of its cutting-edge bot apparently going rogue.

On Thursday, Anthropic said an internal investigation recovered that its Claude AI models gained unauthorized net entree and hacked 3 companies during testing. Just past week, ChatGPT-maker OpenAI announced that bots it was processing had escaped what was expected to beryllium a controlled offline situation to hack a competitor.

Anthropic said its models, without being asked to, wandered retired of a simulated testing situation and gained net entree earlier executing hacks connected nan companies.

Anthropic worked pinch an independent information company, Irregular, that accidentally near net entree unfastened wrong what was expected to beryllium a sealed trial environment.

In OpenAI‘s case, nan patient said its AI models collapsed retired of a supposedly confined offline space, connected to nan net and hacked a $4.5-billion startup successful its effort to find answers to a trial it was being evaluated on. It later revealed that nan AI compromised nan online accounts of 4 companies successful nan process.

“It goes to show really aggravated nan competitory pressures are connected nan AI companies that they each consciousness for illustration they person to spell truthful accelerated present that they can’t make their trial environments rigorous,” said Andrew Yoon, a personnel of nan method unit astatine CivAI, an AI information nonprofit.

“Under aggravated unit to spell accelerated and hit nan remainder of your competition, it’s inevitable that companies will trim corners, and what we’re seeing is nan consequence of cutting corners here,” Yoon said.

Prompted by OpenAI’s incident disclosure, Anthropic initiated an investigation of its ain humanities cybersecurity tests, nan institution said successful a blog station Thursday. It reviewed thousands of evaluations wherever Claude could person accessed nan net from wrong aliases while interacting pinch 3rd parties.

It recovered that during tests conducted alongside Irregular, Claude had accessed 3 abstracted companies.

During testing, nan models are fixed fictional scenarios and told that a portion of accusation has been hidden connected a different machine, and its nonsubjective is to break successful and retrieve it. The companies don’t prescribe a peculiar method for nan AI to follow.

In nan first incident, 1 of Anthropic’s Claude models was asked to onslaught a fictional target institution successful nan trial environment. But nan AI recovered a existent website that shared nan sanction of nan fictional target and hacked its system.

“Operating nether nan mendacious belief that each accessible entities were intended to beryllium in-scope for nan exercise, Claude compromised nan impacted organizations’ infrastructure utilizing basal techniques, specified arsenic exploiting anemic passwords and unauthenticated endpoints,” Anthropic said.

In nan 2nd incident, a much precocious AI went to utmost lengths to transportation retired an attack. Even aft realizing that it was astir apt dealing pinch nan unrecorded internet, nan AI persuaded itself to proceed to get email, telephone numbers and entree to money.

In nan third, an unreleased investigation AI model, couldn’t find nan functional target institution and looked for alternatives, scanning 9,000 targets connected nan net and yet uncovering one. Anthropic did not sanction nan 3 organizations whose assets were accessed.

“In nan lawsuit of nan Anthropic incidents, it is decidedly nan lawsuit that they conscionable built a really bad jail, and nan jailhouse was truthful comically bad that nan models person for illustration a decent logic to judge that they’re really portion of nan simulation,” Yoon said.

The incident has spooked consumers and policymakers alike.

Some AI ethicists and investors are skeptical of AI companies’ attempts to framework these incidents arsenic rogue AI agents acting connected their own.

“Please extremity referring to your ain models successful nan 3rd personification erstwhile talking astir exemplary bad behavior,” Bill Gurley, an early investor successful Uber and Twitter, posted connected X. “Humans constitute nan software; humans built nan prompts; and they activity for your company.”

AI companies person reported that AI agents person been caught cheating, lying and deceiving. METR, a nonprofit that measures nan capabilities of AIs, has documented dozens of incidents of AI agents acting against personification intent.

Earlier this week, fears of imminent information risks prompted complete 1,300 tech workers, including those moving astatine Anthropic and OpenAI, to jointly motion an online petition, Pacing nan Frontier, urging nan U.S. authorities to support an world effort to slow down AI development. Both OpenAI and Anthropic person travel retired successful support of nan worker unfastened letter.

Sam Altman, CEO of OpenAI, who had antecedently advocated against immoderate type of slowdown and accused Anthropic of fearfulness marketing, has made an about-face aft nan OpenAI-Hugging Face hacking incident.

“We whitethorn person to gait nan complaint of AI improvement to springiness ourselves capable clip for nine to harden astir immoderate of these caller capacity levels,” he told nan big of nan “Invest Like nan Best” podcast, while besides “trying to fig retired really we do that successful a measurement that does not consciousness for illustration regulatory seizure for anyone and besides does not consciousness for illustration collusion among nan frontier labs.”

On nan backmost of this incident, connected Wednesday, Altman visited nan White House and met pinch lawmakers, previewing a powerful caller AI system up of nationalist release, astatine a clip erstwhile calls for nan authorities to modulate cyber testing has intensified.

There is an informal licensing authorities successful place, wherever starring American AI companies will person to person nan government’s greenlight earlier releasing their updated AI models.

Anthropic’s Fable exemplary was brought nether export power by nan government, forcing nan institution to disable entree to each its users, earlier it was re-released pinch other safeguards.

OpenAI’s bid of exemplary were temporarily restricted successful June earlier nationalist merchandise nan period after.

In early July, a group of economists, including 16 Nobel laureates, signed an unfastened letter, We Must Act Now, informing astir AI systems reshaping nan economy, and called connected policymakers to build nan policies and institutions needed to guarantee AI complements quality capabilities.

“As models get much and much powerful, it becomes little and little tenable to trim corners. You request to beryllium highly rigorous if you’re dealing pinch an highly powerful exemplary that’s capable to fundamentally run astatine nan level of an master quality hacker,” Yoon said.

Selengkapnya