Current safety alignment training for Large Language Models (LLMs) exhibits a heavy English-centric bias, creating what researchers term a "safety alignment illusion." When these safety filters fail in non-English languages, the consequences are immediate and user-facing, as voice assistants and dialogue systems may produce stereotype-reinforcing outputs and propagate harmful biases.
The INCLUDE Benchmark
To address this cross-lingual gap, researchers introduced INCLUDE (Indian Cultural Lens for Understanding and Detecting Embedded Biases), a multilingual evaluation benchmark. Designed to quantify Indian-centric socio-cultural biases, INCLUDE consists of 2,604 prompts across six languages: English, Hindi, Bengali, Marathi, Tamil, and Hinglish (a Hindi-English code-mix).
Research Findings
The study evaluated ten open- and closed-source LLMs, analyzing a total of 14,988 bias scores. The statistical results revealed two key findings:
- Bengali yielded the highest average bias score in open-source models.
- English demonstrated a notable reversal, producing the lowest bias in open-source models but the highest bias in closed-source models.
These findings highlight a critical failure mode for language technologies deployed across India's linguistically diverse population, where standard English-focused safety alignments are frequently bypassed, allowing cultural biases to persist in non-English outputs.