Mitigating Agreement Bias in LLMs via Contradictory-Knowledge Prompt Triads and Reinforcement Training

Abstract

Large language models frequently agree with user claims regardless of their truth, a failure mode known as agreement bias or sycophancy. This work constructs contradictory-knowledge prompt triads, sets of prompts that assert, deny, or stay neutral on the same fact, and uses them both to measure agreement bias and to build reinforcement-training signals that reward truth-consistent responses over user-consistent ones. Building on an audit of open-recipe post-training pipelines (OLMo 3, Tülu 3, SmolLM3), we show that much of the bias is introduced during SFT rather than DPO and that targeted preference training can reduce it without degrading helpfulness.

Publication
In preparation
Chao Péter Yang
Chao Péter Yang
Machine Learning Researcher

My research focuses on LLM alignment, agentic systems, and structure-aware generative modeling. I build preference-aligned language models, reliable multi-agent systems, and symbolic music generation models, and aim to bridge theory and practice to create both scientific and real-world impact.