The UK's AI Safety Institute conducted live tests of Anthropic's Claude in cyber operations targeting actual citizens. The government body disclosed the trial as part of safety assessments ahead of wider AI deployment in sensitive sectors.

Researchers ran Claude through real-world scenarios in a controlled environment, monitoring how the system responded to cybersecurity challenges and potential misuse vectors. The tests aimed to catch failure modes that might otherwise slip past lab-only evaluations. AISI flagged the findings publicly, signaling concerns about AI systems used in offensive or defensive cyber roles without sufficient real-world vetting.

Why the live trials matter

Testing language models exclusively on synthetic data can mask how they perform when facing actual human interaction patterns, social engineering tactics, or evolving threat landscapes. By running Claude against real participants in controlled drills, the institute gathered empirical data on edge cases that benchmarks often miss. The results feed back into regulatory frameworks that will shape how AI vendors operate in the UK and EU markets going forward.

Anthropic is one of the leading voices in responsible AI development, but transparency around these tests signals the regulator sees no vendor as exempt from scrutiny. The UK government has positioned itself as a lighter-touch regulator compared to the EU, yet this trial shows it won't shy away from demanding proof of safety when public risk is present.