Models from Alibaba, DeepSeek and Moonshot AI frequently repeated official Chinese positions, deflected or refused to answer politically sensitive questions in a benchmark created by Aleph Alpha. Its automated grader rated only 17% to 41% of responses as balanced.

The test used 967 hand-picked prompts covering subjects including Tiananmen, Taiwan and Xinjiang. DeepSeek V4 Pro refused about two-thirds of the questions, while the other tested Chinese models more often returned answers aligned with state doctrine. On ordinary, nonpolitical questions, the measured bias largely diminished but did not disappear completely.

The result is consistent with Chinese rules requiring public-facing AI services to reflect “socialist core values,” as well as earlier audits. However, the methodology deserves scrutiny. Aleph Alpha selected the topics and used its own AI scoring system, so independent reproduction and human review of borderline answers would strengthen the finding.

There is also a commercial conflict: Aleph Alpha sells “sovereign AI” systems to governments and benefits from distinguishing its products from Chinese competitors. The benchmark is useful evidence of uneven behavior on sensitive subjects, but its percentages should be read as one vendor’s measured result under a specific prompt set and grading method—not a universal score for every deployment.