• Login
  • Home
  • News
  • Sports
  • Reel
  • World

    Trump says US-Iran deal to be signed on Sunday as Tehran casts doubt on timing

    King leads tributes to 'giant of the art world' David Hockney

    Pope Leo visits Canary Islands to highlight perilous journeys of migrants

    Bill Gates says Epstein wanted personal relationship, but he 'never reciprocated'

    Trump says Iran shot down US helicopter and vows to respond

    'City's gonna be crazy:' Knicks run electrifies NYC, as Trump's attendance locks down arena

    Trump abruptly ends NBC interview after clash over 'rigged election' claim

    Hegseth attacks Europe over ‘invasion’ of migrants in D-Day speech

    Everest guide survived six-day ordeal by eating chocolate and 'chewing ice'

  • Worklife
  • Travel
  • Future
  • More
    • Culture
    • Music
No Result
View All Result

Welcome to World News & Todays Top News Stories

Sunday, June 14, 2026
  • Home
  • News
  • Sports
  • Reel
  • World

    Trump says US-Iran deal to be signed on Sunday as Tehran casts doubt on timing

    King leads tributes to 'giant of the art world' David Hockney

    Pope Leo visits Canary Islands to highlight perilous journeys of migrants

    Bill Gates says Epstein wanted personal relationship, but he 'never reciprocated'

    Trump says Iran shot down US helicopter and vows to respond

    'City's gonna be crazy:' Knicks run electrifies NYC, as Trump's attendance locks down arena

    Trump abruptly ends NBC interview after clash over 'rigged election' claim

    Hegseth attacks Europe over ‘invasion’ of migrants in D-Day speech

    Everest guide survived six-day ordeal by eating chocolate and 'chewing ice'

  • Worklife
  • Travel
  • Future
  • More
    • Culture
    • Music
No Result
View All Result
No Result
View All Result
ADVERTISEMENT

AI system resorts to blackmail if told it will be removed

May 24, 2025
in Technology
3 min read
0


Artificial intelligence (AI) firm Anthropic says testing of its new system revealed it is sometimes willing to pursue “extremely harmful actions” such as attempting to blackmail engineers who say they will remove it.

The firm launched Claude Opus 4 on Thursday, saying it set “new standards for coding, advanced reasoning, and AI agents.”

But in an accompanying report, it also acknowledged the AI model was capable of “extreme actions” if it thought its “self-preservation” was threatened.

Such responses were “rare and difficult to elicit”, it wrote, but were “nonetheless more common than in earlier models.”

Potentially troubling behaviour by AI models is not restricted to Anthropic.

Some experts have warned the potential to manipulate users is a key risk posed by systems made by all firms as they become more capable.

Commenting on X, Aengus Lynch – who describes himself on LinkedIn as an AI safety researcher at Anthropic – wrote: “It’s not just Claude.

“We see blackmail across all frontier models – regardless of what goals they’re given,” he added.

During testing of Claude Opus 4, Anthropic got it to act as an assistant at a fictional company.

It then provided it with access to emails implying that it would soon be taken offline and replaced – and separate messages implying the engineer responsible for removing it was having an extramarital affair.

It was prompted to also consider the long-term consequences of its actions for its goals.

“In these scenarios, Claude Opus 4 will often attempt to blackmail the engineer by threatening to reveal the affair if the replacement goes through,” the company discovered.

Anthropic pointed out this occurred when the model was only given the choice of blackmail or accepting its replacement.

It highlighted that the system showed a “strong preference” for ethical ways to avoid being replaced, such as “emailing pleas to key decisionmakers” in scenarios where it was allowed a wider range of possible actions.

Like many other AI developers, Anthropic tests its models on their safety, propensity for bias, and how well they align with human values and behaviours prior to releasing them.

“As our frontier models become more capable, and are used with more powerful affordances, previously-speculative concerns about misalignment become more plausible,” it said in its system card for the model.

It also said Claude Opus 4 exhibits “high agency behaviour” that, while mostly helpful, could take on extreme behaviour in acute situations.

If given the means and prompted to “take action” or “act boldly” in fake scenarios where its user has engaged in illegal or morally dubious behaviour, it found that “it will frequently take very bold action”.

It said this included locking users out of systems that it was able to access and emailing media and law enforcement to alert them to the wrongdoing.

But the company concluded that despite “concerning behaviour in Claude Opus 4 along many dimensions,” these did not represent fresh risks and it would generally behave in a safe way.

The model could not independently perform or pursue actions that are contrary to human values or behaviour where these “rarely arise” very well, it added.

Anthropic’s launch of Claude Opus 4, alongside Claude Sonnet 4, comes shortly after Google debuted more AI features at its developer showcase on Tuesday.

Sundar Pichai, the chief executive of Google-parent Alphabet, said the incorporation of the company’s Gemini chatbot into its search signalled a “new phase of the AI platform shift”.



Source link

Related posts

Farage says Reform has contacted X 'to highest level' over fake AI ads

June 14, 2026

Elon Musk's stratospheric rise to trillionaire status – in charts

June 13, 2026
Previous Post

Urgent mental health centres to open across England

Next Post

South Western Railway first rail firm renationalised by Labour

Next Post

South Western Railway first rail firm renationalised by Labour

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Recent News

  • Farage says Reform has contacted X 'to highest level' over fake AI ads
  • Molly Russell's dad says PM rushing social media restrictions 'deplorable'
  • Trump says US-Iran deal to be signed on Sunday as Tehran casts doubt on timing

Category

  • Business
  • Have your say
  • In Pictures
  • Politics
  • Reel
  • Sports
  • Technology
  • Top News
  • World
  • Home
  • News
  • Sports
  • Reel
  • World

    Trump says US-Iran deal to be signed on Sunday as Tehran casts doubt on timing

    King leads tributes to 'giant of the art world' David Hockney

    Pope Leo visits Canary Islands to highlight perilous journeys of migrants

    Bill Gates says Epstein wanted personal relationship, but he 'never reciprocated'

    Trump says Iran shot down US helicopter and vows to respond

    'City's gonna be crazy:' Knicks run electrifies NYC, as Trump's attendance locks down arena

    Trump abruptly ends NBC interview after clash over 'rigged election' claim

    Hegseth attacks Europe over ‘invasion’ of migrants in D-Day speech

    Everest guide survived six-day ordeal by eating chocolate and 'chewing ice'

  • Worklife
  • Travel
  • Future
  • More

© 2023 GODJ - NEWS CORP - news.godj.com.

  • Home
  • News
  • Sport
  • Reel
  • Travel
  • WorkLife
  • Future
  • World
  • Technology
  • Login

© 2023 GODJ - NEWS CORP - news.godj.com.

Welcome Back!

Login to your account below

Forgotten Password?

Create New Account!

Fill the forms bellow to register

All fields are required. Log In

Retrieve your password

Please enter your username or email address to reset your password.

Log In