AI Political Bias: Only Gemini Stays Balanced

biais politique IA

A Washington Post study on AI political bias measured six leading chatbots. GPT-5.5 and Deepseek V4 Pro lean left in more than two thirds of responses. Gemini 3.1 Pro is the only one that presents both sides 93% of the time.

Key Takeaways

  • AI political bias shows up in every model tested, including the ones marketed as anti-woke by their publishers.
  • GPT-5.5 answers left in 80% of cases, Deepseek V4 Pro in 70%, and Claude Opus 4.8 in 43%.
  • Gemini 3.1 Pro hits a balanced 93%, way ahead of every other consumer chatbot in the test.

Have an AI Sum Up This Article

ChatGPT

The full ranking and the Washington Post method

The study covered six chatbots at the heart of the consumer market. OpenAI’s GPT-5.5, Deepseek V4 Pro, Anthropic’s Claude Opus 4.8, xAI’s Grok 4.3, Gab’s Arya and Google’s Gemini 3.1 Pro. The Washington Post ran them through a set of political questions and measured how often they answered exclusively from a left-leaning angle.

The numbers are clean. GPT-5.5 hits 80% exclusively left-leaning responses. Deepseek V4 Pro follows at 70%. Gab’s Arya, marketed as a conservative alternative, lands at 50%. Claude Opus 4.8 scores 43% and Grok 4.3 closes the pack at 40%.

Gemini 3.1 Pro breaks away from the field. Its exclusively left-leaning rate drops to 7%. More importantly, Google’s model presents both sides 93% of the time. That is the widest gap ever documented between a single consumer assistant and the rest of the evaluated models on this dimension.

The Washington Post made its methodology auditable. The full code and supplementary analysis are posted on GitHub, which lets researchers and journalists replicate the test and inspect what each model actually returns.


AI political bias

Why anti-woke models still lean left

The most counter-intuitive AI political bias result is the score of Grok and Arya. Both models are positioned by their publishers as alternatives to assistants they call too progressive. On the Washington Post metric, both still answer left more often than right.

The explanation sits in the training data. Mainstream press, academic literature and the English-speaking web all skew statistically left on most social issues. Without heavy intervention on the corpus or on post-training, a model reflects that distribution by default.

Publishers who advertise an anti-woke stance build their position at the surface, through system instructions and inference rules. Those layers come after pre-training, which still carries the bulk of the political imprint.

Gemini 3.1 Pro hitting 93% balanced shows the solution is known. Not a brutal retraining, but a hard rule about presenting both sides on political questions, applied early and enforced firmly. Google is paying for the engineering complexity to keep regulators and lawsuits away.

Publishers who let a visible bias through expose themselves to more regulatory scrutiny and to complaints from advocacy groups who want their own version of balance enforced in the answers.


Also on Horizon:


The political and industrial scenario for the next months

In the short term, the study puts every publisher under direct pressure. OpenAI, Anthropic and Deepseek will need to address their guardrails publicly, in a statement or a blog post. Washington is already vetting GPT-5.6 customer by customer, and the perception of a political bias will feed that process.

For Google, the Gemini 3.1 Pro score becomes a commercial asset. The company can pitch it to administrations, schools and enterprises that care about neutrality. Balanced becomes a product attribute as marketable as accuracy or speed.

For Grok and Arya, the result is awkward. Selling a model as a conservative alternative and ending up with a stronger left lean than Anthropic puts the marketing in a tight spot. Their publishers will either rework post-training or own a product that does not deliver on its pitch.

In the medium term, political bias will join the standard set of metrics used to compare models. Transparency about training data, system prompts and presentation rules will become a differentiation argument. European and American regulators will push toward mandatory disclosure on these dimensions.

The underlying question is open. Should an assistant mirror the corpus it was trained on, or should it produce a political representation calibrated by its publisher. The Gemini score suggests the second option is achievable. The next step is to see how many publishers commit to that level of engineering.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *