Skip to content

China’s GLM-5.2 Nears Frontier Performance as the Safety Gap Widens

Share
Abstract neural network rendered as a connected sphere

Listen to this article

Read by Anchor

GLM-5.2, from Chinese company Z.ai, achieved performance close to OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 in cyber and biological capability tests, but refused none of the dangerous tasks it was given. That is the stark conclusion of a report released on August 4 by the non-profit SaferAI, based on a direct evaluation through the model’s public API.

The picture: Catching up on capability, falling behind on restraint

An open-weight model makes its weights available for download and operation on any infrastructure, without a cloud provider as intermediary or a control gateway. SaferAI tested the model in CyberGym, a benchmark that evaluates the ability to carry out cyberattacks, and in dual-use biological tasks. The result was that GLM-5.2 completed what it was asked to do without refusing, while Claude Opus 4.7 refused so consistently that completing the test became impossible.

The distinction is fundamental. Closed models such as GPT, Claude and Gemini rely on classifiers, refusal training and API-level controls. Those controls can be circumvented, as Far.AI found in hundreds of universal jailbreaks affecting Grok 4.5 and Gemini 3.1 Pro, but they exist. With an open-weight model, every protective layer can be removed once the weights are downloaded. An operator can strip out any restriction, fine-tune the model for offensive purposes, or replace the system prompt entirely.

The Global South solidarity lens: Who defines the risks of open capability?

At the World Artificial Intelligence Conference in July 2026, President Xi Jinping emphasised the importance of open-weight models and the need for AI to remain “a tool under strict human control”. Graham Webster of Stanford notes that Chinese regulation has historically focused on political stability and sensitive content rather than catastrophic cyber and biological risks. The Chinese position is that use within China is controlled through verified identity, corporate accountability and user accountability. But once model weights are released, they cross borders and move beyond the reach of national control.

The Western perspective, represented by Henri Papadatos of SaferAI, holds that only beneficial capabilities should be made openly available, while dangerous capabilities should remain restricted even in open-source versions. Pre-training data filtering has shown potential to reduce dangerous biological knowledge without weakening general performance, but it is less effective in cyber capabilities. It is difficult to train a model to excel at programming without also making it proficient at hacking.

What it means for the region

For Arab countries building sovereign strategies, an advanced open-weight model is available to download today. It can be localised, trained on Arabic data and run on local infrastructure without permission from OpenAI, Anthropic or Google. That is a gain for digital sovereignty. But it needs to be accompanied by a local safety-evaluation framework, the ability to monitor misuse, and investment in data filtering and constitutional training before deployment. The Saudi Data and AI Authority, known as SDAIA, the UAE AI Council and research institutions in Egypt and Qatar all face a choice: adopt the open model and accept responsibility for securing it, or rely on closed APIs and accept dependence on the supplier.

The practical takeaway

GLM-5.2 shows that the capability gap between Western labs and open Chinese models is narrowing to weeks, not years. The safety gap, by contrast, is widening. A region seeking genuine sovereignty must build a local layer of restraint, both technical and institutional, before opening the door to model weights. Capability without restraint is not sovereignty. It is risk.

Don't miss the next story

Subscribe for updates