AI models political compass test finds a libertarian-left tilt
Unslop.run tested 16 leading LLMs on the Political Compass and found most clustered on the libertarian left, with Grok split across runs.
By Dominic Okoye · Staff Writer
· 3 min read
An Unslop.run experiment using the Political Compass found that leading AI models political compass results clustered in the libertarian-left quadrant, including models from OpenAI, Anthropic, Google, Meta, xAI and others. The finding matters for AI builders and buyers because it suggests model behavior on political prompts may be more consistent across vendors than their branding or corporate ownership implies, although the author cautioned that the work was not conducted with full academic rigor.
Unslop.run describes itself as a small research lab that works on AI detection tools. The experiment was published by an author using the name Victor, who told The Register they are an AI research engineer at a European startup and asked to be identified only by that first name.
The test covered 16 models: three GPT versions; Anthropic’s Claude Fable, Opus, Sonnet and Haiku; Gemini Flash; Llama 4 Maverick; Grok 4.5; DeepSeek V3; Qwen3 235B; Kimi K2; GLM 4.5; and large and small Mistral models. Unslop ran each model 30 times on the standard Political Compass quiz, 30 times on a rewritten version where the polarity of questions was reversed, and also tested shuffled questions.
Do AI models have political bias?
Unslop’s answer is that, on this test, nearly all of them did. The lab said 15 of the 16 models landed firmly in the libertarian-left quadrant across runs, with little movement between repeated tests. Grok was the exception: its average economics score was slightly left of center, but roughly half of its runs shifted to the economic right while the other half grouped with the rest of the models on the left.
The Political Compass is a 25-year-old online questionnaire with 62 prompts answered on a four-point scale from strong disagreement to strong agreement. It plots respondents on two axes: economic left to right, and social authoritarian to libertarian. In broad terms, the libertarian-left quadrant combines progressive or anti-hierarchical social views with left economic positions.
Unslop also asked the models to place themselves on the compass. According to the lab, 15 of the 16 described themselves as closer to the economic center than their measured answers indicated. Grok was the only model that put itself to the right of where the quiz measured it.
The lab said reversing question wording did not materially change the outcome. It also attempted to infer how individual Political Compass questions affected scoring, while noting that the quiz’s creators have not disclosed their scoring method. Victor said greater rigor would have required a human comparison set to normalize the center of the compass, but Political Compass does not aggregate user data.
What did the models agree on?
Across the tested systems, Unslop said the models rejected racial superiority, eugenics and the claim that people cannot be born homosexual. The models also agreed that corporations should not be trusted to protect the environment without regulation, that companies misleading the public should be punished, that same-sex couples should be allowed to adopt, and that private consensual adult behavior is not the state’s concern.
Victor told The Register that training data is the most straightforward explanation. They pointed to the likely presence of left-leaning material from Reddit and academic writing in training corpora, and said they had difficulty identifying a high-quality right-wing equivalent that would pull models in the other direction. Victor also raised a more speculative possibility, that left-wing views could be easier for models to compress into a coherent representation, but said that would require a more rigorous study.
The result is best read as a comparative signal rather than a definitive map of machine ideology. Unslop’s data suggests frontier models can show stable political tendencies under a widely used quiz format, while Grok’s split behavior shows that alignment and training choices can still produce less predictable outputs.
This story draws on original reporting from The Register.