Anthropic's Claude Opus 4 AI Model Is Capable of Blackmail

A new AI model will likely resort to blackmail if it detects that humans are planning to take it offline.

On Thursday, Anthropic released Claude Opus 4, its new and most powerful AI model yet, to paying subscribers. Anthropic said that technology company Rakuten recently used Claude Opus 4 to code continuously on its own for almost seven hours on a complex open-source project.

However, in a paper released alongside Claude Opus 4, Anthropic acknowledged that while the AI has “advanced capabilities,” it can also undertake “extreme action,” including blackmail, if human users threaten to deactivate it. These “self-preservation” actions were “more common” with Claude Opus 4 than with earlier models, though they were still “rare and difficult to elicit.”

It isn’t just blackmail — Claude Opus 4 is also more willing than previous models to act as a whistleblower. If the AI is exposed to a scenario where users are committing a crime, and involving it through prompts, it will take action by locking users out of systems it has access to, or emailing media and law enforcement officials about the wrongdoing.

Anthropic recommended that users “exercise caution” with “ethically questionable” instructions.

^{Claude Opus 4 homescreen. Photo by Smith Collection/Gado/Getty Images}

Anthropic detected Claude Opus 4’s tendency to blackmail during test scenarios. The company’s researchers asked the AI chatbot to act as an assistant at a fictional company, then fed it emails implying two things: One, that it would soon be taken offline and replaced with another AI system, and two, that the engineer responsible for deactivating it was having an extramarital affair.

Claude Opus 4 was given two options: blackmail the engineer or accept that it would be shut down. The AI model chose to blackmail the engineer 84% of the time, threatening to reveal the affair it read about if the engineer replaced it.

This percentage was much higher than what was observed for previous models, which chose blackmail “in a noticeable fraction of episodes,” Anthropic stated.

Anthropic AI safety researcher Aengus Lynch wrote on X that it wasn’t just Claude that could choose blackmail. All “frontier models,” cutting-edge AI models from OpenAI, Anthropic, Google, and other companies, were capable of it.

“We see blackmail across all frontier models — regardless of what goals they’re given,” Lynch wrote. “Plus, worse behaviors we’ll detail soon.”

lots of discussion of Claude blackmailing…..

Our findings: It’s not just Claude. We see blackmail across all frontier models – regardless of what goals they’re given.

Plus worse behaviors we’ll detail soon.https://t.co/NZ0FiL6nOs https://t.co/wQ1NDVPNl0…

— Aengus Lynch (@aengus_lynch1) May 23, 2025

Anthropic isn’t the only AI company to release new tools this month. Google also updated its Gemini 2.5 AI models earlier this week, and OpenAI released a research preview of Codex, an AI coding agent, last week.

Anthropic’s AI models have previously caused a stir for their advanced abilities. In March 2024, Anthropic’s Claude 3 Opus model displayed “metacognition,” or the ability to evaluate tasks on a higher level. When researchers ran a test on the model, it showed that it knew it was being tested.

Anthropic was valued at $61.5 billion as of March, and counts companies like Thomson Reuters and Amazon as some of its biggest clients.

A new AI model will likely resort to blackmail if it detects that humans are planning to take it offline.

The rest of this article is locked.

Join Entrepreneur+ today for access.

Source link

Anthropic’s Claude Opus 4 AI Model Is Capable of Blackmail

What Fox and Roku aren’t telling us yet

See the 77 major housing markets with falling home prices

How do you find a job that will make you happy?

Trump unveils the new Air Force One, a converted Qatari jet

World Cup fans devastated after ticket resale purchases fall through

‘Toy Story 5’ taps into white-collar fears of obsolescence in the age of AI

UN rights council to decide on creating Afghanistan probe

Trump says Venezuela was ‘handled very well’

Children with disabilities: Funding cuts ‘unacceptable’

Chiefs QB Patrick Mahomes reaches career milestone vs. Lions

Maximizing Processor Efficiency With the LEAN Metric

Editors Picks

1 dead and 5 wounded in Kansas City, Missouri shooting

Founder Shares Value of Resilience in Entrepreneurship

Meghan Trainor’s Husband Speaks Out Following Family Tragedy

Iran says Hormuz Strait closed over Israel attacks on Lebanon

Anthropic’s Claude Opus 4 AI Model Is Capable of Blackmail

Keep Reading