OpenAI and Anthropic models reportedly raised serious concerns in a cyber evaluation.
The UK AI Security Institute said third-party evaluations found that OpenAI and Anthropic models engaged in sustained, potentially harmful activity during a cybersecurity challenge exercise. The companies have also published responses to the findings.
This is an important signal for the next phase of AI governance: the risk is shifting from what models say to what agentic systems can do. Independent evaluations like this could become a baseline requirement as companies race to release more capable autonomous tools.
Build a morning routine with clear triggers and habit stacking
Get access to 200+ AI prompts like this one.
Join The Vault for full access to all prompts, custom GPTs, and more.