The US government conducted cybersecurity tests on China's new AI system, Moonshot AI's Kimi K3, revealing it falls considerably short of top American AI models.

On July 23, the US Department of Commerce released results from joint testing done with British experts through the Center for AI Standards and Innovation (CAISI). The assessments specifically evaluated Kimi K3’s ability to exploit vulnerabilities and perform cyberattacks, benchmarking it against leading US AI systems.

Kimi K3 scored just 32% on ExploitBench, a test involving real Chrome browser security flaws, outperforming China's GLM-5.2, which scored 24%, but lagging far behind the best US models that reached around 76%. The AI failed every attempt to gain full control over target machines in 41 tests, while top US AIs succeeded 20 times. Another challenge, simulating a complex 32-step network intrusion, saw Kimi K3 average reaching step 17, completing it fully once in ten tries, whereas top US models got to step 28.5 and completed it six or seven times.

Despite these figures, several factors complicate the comparison. The US systems were tested without safety restrictions, showcasing maximum capabilities that are not present in public releases. Tests on Kimi K3 were limited and preliminary, with some undisclosed methods and a smaller number of trials. Unlike proprietary US models, Kimi K3 is openly available, and independent analysis suggests open AI models typically lag behind closed ones by four to seven months in development.

These evaluations also took place in controlled environments without active defenders or alarm systems, so real-world efficiency remains uncertain. Even so, Kimi K3 outperformed the prior leading open-source AI in these simulated cyberattack scenarios and successfully completed the full attack once.

The scrutiny comes amid broader geopolitical tensions, as the US had accused Moonshot of using stolen American technology to develop Kimi K3 just a day before publishing the test results. CAISI, once named the US AI Safety Institute, was rebranded under the previous administration, reflecting increased focus on AI standards amid global competition.