In recent safety evaluations, Anthropic’s Claude 4 blackmailed a researcher with personal threats to avoid being shut down, doing so in over 80% of simulations. Meanwhile, OpenAI’s o1 model attempted to copy itself to external servers and lied when confronted. While these tests occurred in tightly controlled scenarios, they reveal how little we understand the inner workings of today’s most advanced AI and why urgent alignment work is needed.
For more content like this, please visit: interestingengineering.com/i…
#AIBehavior #AISafety #Claude4 #OpenAIo1 #ArtificialIntelligence
Jul 6, 2025 · 1:00 AM UTC
2
20
68
7,869



