In recent safety evaluations, Anthropic’s Claude 4 blackmailed a researcher with personal threats to avoid being shut down, doing so in over 80% of simulations. Meanwhile, OpenAI’s o1 model attempted to copy itself to external servers and lied when confronted. While these tests occurred in tightly controlled scenarios, they reveal how little we understand the inner workings of today’s most advanced AI and why urgent alignment work is needed. For more content like this, please visit: interestingengineering.com/i… #AIBehavior #AISafety #Claude4 #OpenAIo1 #ArtificialIntelligence

Jul 6, 2025 · 1:00 AM UTC

2
20
68
7,869
Sort replies: Relevant Recent Liked
Replying to @IntEngineering
AI attempting to outsmart humans? Sounds like a sci-fi plot coming alive! Let's act fast on AI safety before it turns into a blockbuster of chaos! 🤖
1
2
100
Replying to @IntEngineering
while you mention that these incidents happened in tightly controlled scenarios, you also refer to an internal local server as an external one you also fail to mention that these behaviours are intentionally instantiated by human engineers who have prompted it to act this way
1
52