The alignment of large language models (LLMs) with human values is a central goal of AI research, but a new study published on arXiv (cs.AI) suggests that current evaluation methods may be misleading. The research, titled "Agreement Is Not Alignment," argues that agreement on final moral judgment labels does not prove that models and humans rely on the same moral grounds.
The Illusion of Moral Convergence
Researchers utilized a curated 500-item benchmark derived from the ETHICS dataset, spanning five domains of moral judgment. The study compared responses from human annotators with those from frontier and open-source LLMs, examining not only the final labels but also the supporting rationales.
While models often showed high agreement with the human majority on final judgments, rationale-level analysis revealed a systematic divergence. Specifically, models redistribute attention across various moral categories, such as:
- Harm
- Respect
- Promise-keeping
- Justice
- Desert
- Excuse relevance
This redistribution occurs even when the model's final label matches the human consensus, indicating a different underlying logic.
Risks of Label-Based Evaluation
According to the study, evaluations based solely on final labels can provide "misleading reassurance." Two agents may reach the same conclusion while appealing to different principles, contextual assumptions, or interpretations of a situation.
The results highlight the necessity of complementing label-based metrics with an analysis of the reasons, principles, and moral priorities expressed in model judgments. Without this, the technical community risks assuming alignment where only superficial agreement exists.