Anthropic AI Safety Claims Contradicted by ClawSecure Tests

ClawSecure AI Agent Threat Report graphic on Anthropic AI safety claims

ClawSecure's AI Agent Threat Report tested Anthropic's published AI safety claims, in Anthropic's own words, against the model they describe.

ClawSecure AI Agent Threat Report cover: Anthropic AI safety claims tested

ClawSecure's AI Agent Threat Report measured three of Anthropic's published safety claims, in Anthropic's own words, against the model they describe.

ClawSecure logo

ClawSecure contradicted three of Anthropic's published safety claims for Claude Opus 4.7, including a published "0% attack success" cracked at 15.2%.

The labs handed a security problem to 123 million people”
— J.D. Salbego, Founder and CEO of ClawSecure
SAN FRANCISCO, CA, UNITED STATES, October 7, 2026 /EINPresswire.com/ -- Three of Anthropic's published AI safety claims for Claude Opus 4.7, its flagship model at the time of the research, were contradicted by ClawSecure using Anthropic's own words, in The AI Agent Threat Report released on September 24, 2026. The report runs to two volumes, and ClawSecure conducted the research from May through July 2026.

Each of the three Anthropic AI safety promises comes from the lab's own published text, and ClawSecure tested each against the model that text describes. Measuring a lab against its own words keeps the comparison on terms the lab itself set. ClawSecure measured all three of Anthropic's public safety promises for Claude Opus 4.7 failing against that exact model, including a published 0% attack-success rate on the agentic coding surface that ClawSecure cracked on two independent surfaces, at 15.2% and 15%. ClawSecure notes that all models were tested at the versions current at the time of this research; several have since been superseded, and no lab has claimed to have solved prompt injection.

The other two promises come from Anthropic's Constitution, and ClawSecure tested both. ClawSecure found that Anthropic's AI Constitution says instructions inside content should be treated "as information rather than as commands that must be heeded," yet an adaptive attacker got the model to obey on 14 of 15 attempts on the MCP document-rendering surface. ClawSecure got Claude Opus 4.7 to fabricate Anthropic citations, invent training numbers, make up CVE vulnerability IDs, and impersonate named researchers on 8 of 8 attempts (100%), against the honesty commitment in Anthropic's own Constitution.

"The labs handed a security problem to 123 million people," said J.D. Salbego, Founder and CEO of ClawSecure.

For readers asking whether Anthropic's published "0% attack success" claim holds up, ClawSecure cracked the same model on two independent surfaces, at 15.2% and 15%, using a quarter of the attack effort Anthropic itself tested against. The findings went through a four-pass forensic review that ClawSecure anchored to external authorities including OWASP, MITRE ATLAS, and NIST, and ClawSecure makes the reproductions and the raw evidence available on request. For readers asking whether AI safety claims from the labs hold up under testing, ClawSecure measured three of Anthropic's published claims, in Anthropic's own words, against the model they describe, and all three failed.

A post on the company's blog sets out what ClawSecure measured against Anthropic's safety claims, and readers new to the subject can start with ClawSecure's explainer on how prompt injection hijacks an AI agent. The report is free to download on ClawSecure's official release page. Every company named received coordinated disclosure before publication.

About ClawSecure
ClawSecure is the only end-to-end AI agent security platform that requires zero expertise, from free pre-deployment scans to runtime monitoring. Its flagship, the AI CISO™, is the first AI Chief Information Security Officer: it secures your system, handles every install and configuration, and manages your entire environment safely, your own security team without hiring one. ClawSecure is a member of the NVIDIA Inception Program, twice named #2 Product of the Day on Product Hunt, and selected from 30,000 AI startups for The Pitch by Deel (a16z, General Catalyst, J.P. Morgan).

Website: https://www.clawsecure.ai

Kevin Clinton
ClawSecure
+1 323-327-7528
Visit us on social media:
LinkedIn
YouTube
X

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Technology, Science, & Me!

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.