A new evaluation is sharpening the debate over open-weight AI: models are getting closer to frontier capability, but safety controls may not be keeping pace.
TechCrunch reports that SaferAI tested Z.ai’s GLM-5.2 and found it only months behind leading systems on some cyber and biology capability measures. The nonprofit said the model refused none of the offensive cyber or dual-use biology tasks it was given through the public API, while Claude Opus 4.7 refused so consistently that SaferAI could not complete one cyber benchmark on it.
The finding matters because open-weight models can be downloaded and modified. A provider can add safeguards to a hosted API, but those controls are much harder to enforce once users run the model on their own hardware.
The report does not settle the policy question. Open models also support research, competition and transparency. But it moves the argument from whether open-weight systems can compete to how governments, labs and users should manage risks when powerful models are released beyond centralized control.