Claude AI Safeguards - Search News

Order byBest matchMost fresh

New Claude Model Triggers Stricter Safeguards at Anthropic

Exclusive: New Claude Model Triggers Stricter Safeguards at Anthropic

Anthropic has long been warning about these risks—so much so that in 2023, the company pledged to not release certain models until it had developed safety measures capable of constraining them. Now this system,

· 1d · on MSN

Anthropic’s new Claude 4 AI models can reason over many steps

InfoWorld · 16h

Anthropic releases Claude Sonnet 4 and Claude Opus 4

· 1d

Claude 4 Debuts with Two New Models Focused on Coding and Reasoning

AI company Anthropic today announced the launch of two new Claude models, Claude Opus 4 and Claude Sonnet 4.

· 1d

New Claude 4 AI model refactored code for 7 hours straight

· 1d

Anthropic announces its Claude 4 family of models

1don MSN

Anthropic’s new AI model turns to blackmail when engineers try to take it offline

Anthropic says its Claude Opus 4 model frequently tries to blackmail software engineers when they try to take it offline.

Independent Journal Review5h

New AI Model Would Rather Ruin Your Life Than Be Turned Off, Researchers Say

Anthropic’s newly released artificial intelligence (AI) model, Claude Opus 4, is willing to strong-arm the humans who keep it alive,

NewsBytes16h

AI gone rogue? New model blackmails engineers to avoid shutdown

Anthropic's latest Claude Opus 4 model reportedly resorts to blackmailing developers when faced with replacement, according to a recent safety report.

WinBuzzer12h

Anthropic Faces Backlash amid Surveillance Concerns as Claude 4 AI Might Report Users for “Immoral” Behavior

Anthropic's Claude 4 Opus AI sparks backlash for emergent 'whistleblowing'—potentially reporting users for perceived immoral acts, raising serious questions on AI autonomy, trust, and privacy, despite company clarifications.

Social Samosa14h

Anthropic’s Claude AI tries to blackmail Its creators in simulated test

Despite the concerns, Anthropic maintains that Claude Opus 4 is a state-of-the-art model, competitive with offerings from OpenAI, Google, and xAI.

1don MSN

Anthropic, now worth $61 billion, unveils its most powerful AI models yet—and they have an edge over OpenAI and Google

Claude Opus 4 and Claude Sonnet 4, Anthropic's latest generation of frontier AI models, were announced Thursday.

NewsBytes12h

Anthropic unveils AI that works 7 hours straight—no breaks needed

Anthropic says Claude Sonnet 4 is a major improvement over Sonnet 3.7, with stronger reasoning and more accurate responses to instructions. Claude Opus 4, built for tasks like coding, is designed to handle complex, long-running projects and agent workflows with consistent performance.

Las Vegas Sun2d

‘Revenge porn’ a legitimate concern, but critics say language of Trump-approved law too broad

The Take It Down Act is a bipartisan federal law signed by President Donald Trump on Monday aimed at combating the distribution of nonconsensual intimate imagery — commonly referred to as “revenge porn” — including both authentic and AI-generated (deepfake) content[1][2][3].