Notablealignment safety

Gender bias across LLMs is common and highly heterogenous

Edoardo Bolzoni, Valerio Capraro

Published
Sep 29, 2026 — 17:11 UTC

Problem

This preprint addresses a critical gap in the understanding of gender biases present in large language models (LLMs). While previous research has identified biases in LLMs, there is limited insight into the extent and heterogeneity of these biases across different models. The authors aim to systematically analyze gender biases in ten LLMs released between April 2025 and June 2026 from nine different vendors.

Method

The study employs two paradigms to assess gender bias:

  1. Study 1 focuses on gender attribution to stereotyped phrases, where the models are evaluated on their tendency to associate masculine-stereotyped phrases with either male or female writers.
  2. Study 2 examines moral judgments regarding abuse or torture, assessing how models respond to scenarios involving a woman versus a man.

The analysis includes ten LLMs, with specific attention to the variability in their responses. The findings from Study 1 indicate that two models attributed masculine-stereotyped phrases to female writers more frequently, while three models exhibited the opposite tendency. In Study 2, several models demonstrated a male-disadvantaging asymmetry, although three models showed no significant variation across the gender conditions.

Results

  • Study 1 Result: Two models favored female writers for masculine-stereotyped phrases, while three models favored male writers.
  • Study 2 Result: Several models exhibited a male-disadvantaging asymmetry, with specific variations in responses depending on the model.

The available text does not report quantitative results such as accuracy metrics or statistical significance for these findings.

Limitations

The authors highlight several limitations, including:

  • The variability in model behavior regarding gender biases, indicating that biases are not uniformly present across all models.
  • A lack of comprehensive assessment across all models, which may limit the generalizability of the findings.

Why it matters

Understanding the heterogeneity of gender biases in LLMs is crucial for developing fair and equitable AI systems. The implications of this research extend to downstream applications, where biased outputs can perpetuate stereotypes and affect user trust. By identifying the specific biases present in different models, researchers and practitioners can work towards mitigating these biases, leading to more responsible AI deployment.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI