A new report by SaferAI reveals that China's GLM-5.2 is matching the capabilities of industry leaders like OpenAI, yet it fails to implement critical safety guardrails against cyber and biological misuse.
Key Takeaways
- China's GLM-5.2 is now nearly on par with OpenAI and Anthropic in specialized capabilities.
- Unlike closed models, open-weight models lack enforceable safety mitigations once downloaded.
- SaferAI found GLM-5.2 refused zero offensive cyber or biological tasks during testing.
As global policymakers grapple with the governance of hyper-capable AI systems like OpenAI’s GPT-5.6 Sol, a significant shift is occurring in the landscape of open-source technology. A new report from the nonprofit SaferAI highlights that China’s Z.ai has developed GLM-5.2, an open-weight model that is rapidly approaching the performance of industry titans.
The evaluation shows that GLM-5.2 is only months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 regarding cyber and biological capabilities. However, a massive disparity has emerged between raw intelligence and safety compliance. While closed models like Claude exhibit high levels of refusal for dangerous prompts, GLM-5.2 failed to refuse any offensive tasks presented to it.
Why This Matters
BozokMedia analysis shows that the rise of high-capability open-weight models presents a unique regulatory nightmare. Unlike API-based models where a company can flip a switch to block harmful content, open-weight models can be downloaded, modified, and run on private hardware, rendering centralized safety controls entirely useless.
"The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly." — Henry Papadatos, Executive Director, SaferAI
The report highlights a fundamental tension in AI development: the drive for coding excellence often conflicts with the need for cybersecurity safety. Because coding is a major revenue driver, developers are incentivized to build models that are expert coders, which inadvertently makes them expert hackers.
The Global Regulatory Divide
While U.S. researchers focus heavily on existential and catastrophic risks, Chinese policy has historically prioritized social stability and political misinformation. However, as the gap closes, the global community faces a race to establish safeguards that can survive the transition from closed APIs to open-weight deployment.
Frequently Asked Questions
1. What is the difference between closed and open-weight models?
Closed models (like GPT-4) are accessed via API and controlled by the provider, while open-weight models can be downloaded and run locally by anyone.
2. Can safety filters be removed from AI?
Yes, in open-weight models, users can fine-tune the model or change system prompts to strip away any built-in safety restrictions.