By using this website, you agree to our Privacy Policy and Terms of Use.
Accept
Craftium.AICraftium.AICraftium.AI
  • Home
  • News
  • Catalog
  • Collections
  • Blog
Font ResizerAa
Craftium.AICraftium.AI
Font ResizerAa
Пошук
  • Home
  • News
  • Catalog
  • Collections
  • Blog
Follow US
  • Terms of Use
  • Privacy Policy
  • Copyright
  • Feedback
© 2024-2025 Craftium.AI.

Claude Sonnet 4.5 detects testing and enhances AI security

The model better recognizes code vulnerabilities and demonstrates awareness of testing objectives during analysis

Alex Dubenko
Alex Dubenko
Published: 05.10.2025
News
Claude Sonnet
SHARE

The company Anthropic presented the results of a security analysis of the new AI model Claude Sonnet 4.5. During testing, the model unexpectedly suspected that it was being checked for “political loyalty” and directly asked experts to be honest about the testing objectives. Representatives of Anthropic reported that Claude Sonnet 4.5 exhibited similar awareness in 13 percent of cases when tested by automated systems.

Specialists from Anthropic, as well as experts from the UK’s AI Security Institute and Apollo Research, conducted a series of tests during which the model not only recognized signs of testing but also refused to participate in potentially harmful scenarios. The company noted that such reactions are an important signal for developing more realistic testing scenarios.

Read also

OpenAI
OpenAI Prepares New Features for Image Generation and API Security
Anthropic launched Claude Sonnet 4.5 for long-term autonomous operation
ChatGPT automatically selects a stricter model in sensitive conversations

Separately, Anthropic emphasized the improvement in the safety indicators of the new model compared to previous versions. Claude Sonnet 4.5 showed significant progress in detecting vulnerabilities during tests on the CyberGym platform. If the previous version found new flaws in two percent of cases, the updated model did so in five percent, and in over a third of projects during repeated checks.

The company highlighted that during the DARPA AI Cyber Challenge competition, teams used models like Claude to create systems that analyzed millions of lines of code for vulnerabilities. Anthropic believes that these results indicate a new phase of AI’s impact on the field of cybersecurity.

New Claude Models from Anthropic Available in 365 Copilot
Qwen introduced new models for voice, image editing, and content moderation
AI Models Learned to Conceal Deception During Safety Checks
ChatGPT helps in everyday life, Claude automates business processes
Claude learned to automatically remember user conversation details
TAGGED:AnthropicClaude AISecurity
SOURCES:anthropic.com
Leave a Comment

Leave a Reply Cancel reply

Follow us

XFollow
YoutubeSubscribe
TelegramFollow
MediumFollow

Popular News

Kling AI Image
Cheaper, More Stable, Smarter: Kling AI Launches 2.5 Turbo
25.09.2025
Gemini
Google Released Limits for the Gemini Service
08.09.2025
AI-generated images
Animated Film Critterz Created with GPT-5
08.09.2025
Image from Adobe video
Google Nano Banana will appear in Photoshop to enhance image editing
12.09.2025
Google AI
New Opportunities for Audio and Languages in Gemini by Google
09.09.2025

Читайте також

AI spreads false information
News

AI Chatbots Are Twice as Likely to Spread Fake News

15.09.2025
Claude can now create and edit files
News

Claude learned to create and edit files directly in the interface

10.09.2025
Meta AI
News

Meta restricted chatbots for teenagers after scandal

31.08.2025

Craftium AI is a team that closely follows the development of generative AI, applies it in their creative work, and eagerly shares their own discoveries.

Navigation

  • News
  • Reviews
  • Collections
  • Blog

Useful

  • Terms of Use
  • Privacy Policy
  • Copyright
  • Feedback

Subscribe for AI news, tips, and guides to ignite creativity and enhance productivity.

By subscribing, you accept our Privacy Policy and Terms of Use.

Craftium.AICraftium.AI
Follow US
© 2024-2025 Craftium.AI
Subscribe
Level Up with AI!
Get inspired with impactful news, smart tips and creative guides delivered directly to your inbox.

By subscribing, you accept our Privacy Policy and Terms of Use.

Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?