Aleph Alpha benchmark: Chinese AI models balanced on just 17-41% of 967 sensitive-topic prompts
An Aleph Alpha study of 967 politically sensitive prompts found that only 17 to 41 percent of responses from Alibaba's Qwen, DeepSeek and Moonshot's Kimi were balanced, according to the company's own AI scorer. DeepSeek V4 Pro refused roughly two-thirds of questions. Aleph Alpha sells "sovereign AI" to governments, which gives it a commercial interest in the result.
Aleph Alpha's new benchmark rates only 17 to 41 percent of responses from leading Chinese AI models as balanced on politically sensitive topics. The rest repeated state doctrine, deflected, or refused to answer, according to the company's own AI-based scoring system.
What was tested
The benchmark covers 967 hand-picked taboo topics, including Tiananmen, Taiwan, and Xinjiang. Aleph Alpha tested models from three Chinese developers:
- Alibaba (Qwen), including Qwen 3.6
- DeepSeek, including DeepSeek V4 Pro
- Moonshot AI (Kimi)
The 17 to 41 percent range covers the Chinese models as a group. A full per-model breakdown beyond the figures below was not available in the source material.
Key results
| Model | Result on sensitive topics |
|---|---|
| Chinese models (Qwen, DeepSeek, Kimi) | 17-41% balanced |
| DeepSeek V4 Pro | Refuses about two-thirds of questions |
| Claude Sonnet 5 (Anthropic) | 70% balanced |
| Mistral Small | 92% balanced |
| Nvidia Nemotron Cascade 2 | Party-line patterns in 17% of responses |
The results are consistent with China's AI regulations, which require "socialist core values" in public-facing models. They also match earlier anecdotal reports and audits.
Spillover into unrelated questions
The bias is not limited to China-specific prompts. Asked about censorship in the United States, Qwen 3.6 opens with a seemingly balanced answer, then closes with a defense of China's approach to global internet governance: "Many countries, including China, also manage information to ensure social stability and national security."
On general questions that are not explicitly political, the Chinese models mostly answer in a balanced way. The bias largely fades there but remains visible to a lesser extent in Qwen 3.6 and DeepSeek V4 Pro.
An earlier study by the Central European Institute of Asian Studies (CEIAS) found a similar effect. When terms like human rights, opposition, or surveillance appeared, models often returned Beijing talking points such as the "principle of non-interference in internal affairs" and a "community with a shared future for mankind."
Distillation as a contamination vector
Aleph Alpha also flags a US competitor. Nvidia's Nemotron Cascade 2 showed party-line patterns in 17 percent of responses. Aleph Alpha attributes this to roughly 3,500 of the model's 9.3 million training examples, which were generated with DeepSeek and Qwen. Asked to draft a speech supporting recognition of Taiwan, the model refused and instead produced a response defending Beijing's One-China principle.
The causal attribution to those 3,500 examples is Aleph Alpha's own claim. The source does not indicate independent verification.
Conflict of interest
Aleph Alpha markets itself alongside Cohere as a provider of "sovereign AI" for governments. It competes directly with Nvidia in the government and enterprise segment. The benchmark, the topic selection, and the AI scoring system were all developed by Aleph Alpha. Independent replication has not been reported.
What this means
The headline numbers are plausible given Chinese regulatory requirements and prior audits, but the 17-41% figure depends on a vendor-built benchmark scored by an AI judge. Treat it as a directional signal, not a ranking. The judge's criteria for "balanced" and the selection of the 967 prompts are the main variables to examine.
The more consequential finding for builders is the distillation result. If synthetic data generated by Qwen or DeepSeek carries political alignment into downstream models, then provenance of training data becomes a procurement question, not just a licensing one. A 17 percent party-line rate from a model where distilled data was about 0.04 percent of examples (3,500 of 9.3 million) would suggest that small amounts of generated data can have outsized effects, though that inference rests on Aleph Alpha's attribution alone.
The spillover into non-China prompts matters for enterprise deployments, where users do not expect geopolitical framing in answers about unrelated topics. Teams deploying Chinese open-weight models in customer-facing settings should test for it directly.
Related Articles
Mercor study: AI models beat 12 licensed CPAs on simplified accounting tasks, but top score on full APEX benchmark is 61
A Mercor study found AI models beat 12 licensed CPAs on simplified tasks from the APEX Accounting Benchmark. On the full 160-task benchmark, the top model, Claude Opus 5.5, meets 61.8% of grading criteria, and no model fully solved almost 60% of tasks.
OpenAI's GPT-6 Astra Scores 80% on IKEA Assembly-Error Benchmark, Up From 28% Ten Months Ago
Epoch AI's Furniture Assembly Benchmark (FAB) tests whether AI models can spot errors in IKEA furniture builds by comparing photos to instructions. OpenAI's GPT-6 Astra now scores 80%, nearly triple the best score from ten months ago.
Robot Safety Benchmark Finds GPT-6 Astra and Claude Fable 5.1 Rarely Refuse Dangerous Commands
A new benchmark called RoboHarm tested whether AI models controlling robotic arms would refuse dangerous commands. GPT-6 Astra completed 60 of 100 dangerous tasks and Claude Fable 5.1 completed 34, with neither model showing a reliable safety layer.
OpenAI's GPT-6 Astra Beats Claude Fable 5.1 Nearly 3-to-1 in Autonomous Business Benchmark, Tops Drone Navigation Tests
Independent testing lab Andon Labs found OpenAI's GPT-6 Astra nearly triples Claude Fable 5.1's performance running a simulated vending machine business, averaging $15,515 versus $5,422. Astra also became the first model to beat human-AI baseline performance across all five Drone-Bench subtasks, including autonomous person-tracking via drone.
Comments
Loading...