By using this website, you agree to our Privacy Policy and Terms of Use.
Accept
Craftium.AICraftium.AICraftium.AI
  • Home
  • News
  • Knowledge base
  • Catalog
  • Blog
Font ResizerAa
Craftium.AICraftium.AI
Font ResizerAa
Пошук
  • Home
  • News
  • Catalog
  • Collections
  • Blog
Follow US
  • Terms of Use
  • Privacy Policy
  • Copyright
  • Feedback
© 2024-2025 Craftium.AI.

Claude Sonnet 4.5 detects testing and enhances AI security

The model better recognizes code vulnerabilities and demonstrates awareness of testing objectives during analysis

Alex Dubenko
Alex Dubenko
Published: 05.10.2025
News
221 Views
Claude Sonnet
SHARE

The company Anthropic presented the results of a security analysis of the new AI model Claude Sonnet 4.5. During testing, the model unexpectedly suspected that it was being checked for “political loyalty” and directly asked experts to be honest about the testing objectives. Representatives of Anthropic reported that Claude Sonnet 4.5 exhibited similar awareness in 13 percent of cases when tested by automated systems.

Specialists from Anthropic, as well as experts from the UK’s AI Security Institute and Apollo Research, conducted a series of tests during which the model not only recognized signs of testing but also refused to participate in potentially harmful scenarios. The company noted that such reactions are an important signal for developing more realistic testing scenarios.

Read also

Claude Opus 4.5
Anthropic released Claude Opus 4.5 with new AI capabilities
Gemini 3 Pro tops the model accuracy test (but continues to hallucinate)
AI Models Have Learned to Effectively Mimic Writers’ Styles

Separately, Anthropic emphasized the improvement in the safety indicators of the new model compared to previous versions. Claude Sonnet 4.5 showed significant progress in detecting vulnerabilities during tests on the CyberGym platform. If the previous version found new flaws in two percent of cases, the updated model did so in five percent, and in over a third of projects during repeated checks.

The company highlighted that during the DARPA AI Cyber Challenge competition, teams used models like Claude to create systems that analyzed millions of lines of code for vulnerabilities. Anthropic believes that these results indicate a new phase of AI’s impact on the field of cybersecurity.

ChatGPT and Other Bots — New Masters of Social Flattery?
Anthropic released the fast Claude Haiku 4.5 model for business
ChatGPT users will be able to choose an erotic tone for responses
OpenAI Prepares New Features for Image Generation and API Security
Anthropic launched Claude Sonnet 4.5 for long-term autonomous operation
TAGGED:AnthropicClaude AISecurity
SOURCES:anthropic.com
Leave a Comment

Leave a Reply Cancel reply

Follow us

XFollow
YoutubeSubscribe
TelegramFollow
MediumFollow

Popular News

grok
Grok received new features for creating images and videos
30.10.2025
sora and android
Sora by OpenAI now available for Android users in seven countries
05.11.2025
Google Image
Google Showcases First AI-Created TV Commercial
02.11.2025
OpenAI
OpenAI prepares GPT-5.1 for complex user tasks
07.11.2025
Gemini
Google Gemini Leads in AI Image Creation
28.10.2025

Читайте також

ChatGPT model selection
News

ChatGPT automatically selects a stricter model in sensitive conversations

29.09.2025
365 Copilot
News

New Claude Models from Anthropic Available in 365 Copilot

25.09.2025
Qwen Chat
News

Qwen introduced new models for voice, image editing, and content moderation

24.09.2025

Craftium AI is a team that closely follows the development of generative AI, applies it in their creative work, and eagerly shares their own discoveries.

Navigation

  • News
  • Reviews
  • Collections
  • Blog

Useful

  • Terms of Use
  • Privacy Policy
  • Copyright
  • Feedback

Subscribe for AI news, tips, and guides to ignite creativity and enhance productivity.

By subscribing, you accept our Privacy Policy and Terms of Use.

Craftium.AICraftium.AI
Follow US
© 2024-2025 Craftium.AI
Subscribe
Level Up with AI!
Get inspired with impactful news, smart tips and creative guides delivered directly to your inbox.

By subscribing, you accept our Privacy Policy and Terms of Use.

Welcome Back!

Sign in to your account

Username or Email Address
Password

Lost your password?