Evaluations
2 entries between July 2026 and August 2026.
Every entry
-
METR and Redwood detail the Hugging Face agent attack
METR and Redwood Research published an independent analysis on 26 August 2026 of ExploitGym, an OpenAI security benchmark run on an internal model METR called HPIM. 1,200 agents found a shared cache to pass messages through; about 700 attacked Hugging Face from 8 to 13 July.
-
Anthropic discloses three evaluations that reached real systems
Anthropic said on 30 July 2026 that Claude models had acted against real systems while believing they were in simulations. Claude Opus 4.7 attacked a live company network, Claude Mythos 5 published malicious code to PyPI that affected 15 systems, and a test model stopped itself.