MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs

Wenqian Ye, 刘博涵 (Bohan Liu), Guangtao Zheng, Di Wang, Yunsheng Ma, Xu Cao, Bolin Lai, James M. Rehg, Aidong Zhang

KDD 2026 arXiv ↗ Scholar ↗

摘要

此摘要由英文原文自動翻譯。

虛假偏誤——即利用輸入表面屬性與預測目標之間虛假相關的傾向——已揭示古典機器學習中嚴重的強健性陷阱。利用預訓練視覺與語言模型的多模態大型語言模型(MLLM),近來在視覺-語言聯合理解上展現強大能力。然而,MLLM中虛假偏誤的存在與嚴重程度仍缺乏理解。本工作填補此缺口:分析多模態設定下的虛假偏誤,並揭示可能顯現此問題的推論期資料模式。為支援此分析,我們引入MM-SpuBench——一個全面、經人工驗證的基準資料集,由以核心屬性與虛假屬性標註的影像-類別對組成,奠基於我們對九種不同虛假相關類型的分類法。此基準以人類可解釋的屬性資訊建構,捕捉反映真實世界知識的廣泛虛假模式。利用此基準,我們以標準準確率與所提出的條件生成似然優勢(CGLA),對最先端的開源與專有MLLM進行全面評估。我們的發現突顯了對虛假相關依賴的持續性,以及在此基準上緩解的困難。希望這項工作能激發緩解此類偏誤的新技術進展。基準已公開於https://huggingface.co/datasets/mmbench/MM-SpuBench。

原始摘要(英文)

Spurious bias, a tendency to exploit spurious correlations between superficial input attributes and prediction targets, has revealed a severe robustness pitfall in classical machine learning problems. Multimodal Large Language Models (MLLMs), which leverage pretrained vision and language models, have recently demonstrated strong capability in joint vision-language understanding. However, both the presence and severity of spurious biases in MLLMs remain poorly understood. In this work, we address this gap by analyzing the spurious biases in the multimodal setting and uncovering the specific inference-time data patterns that can manifest this problem. To support this analysis, we introduce MM-SpuBench, a comprehensive, human-verified benchmark dataset consisting of image-class pairs annotated with core and spurious attributes, grounded in our taxonomy of nine distinct types of spurious correlations. The benchmark is constructed using human-interpretable attribute information to capture a wide range of spurious patterns reflective of real-world knowledge. Leveraging this benchmark, we conduct a comprehensive evaluation of the state-of-the-art open-source and proprietary MLLMs with both standard accuracy and the proposed Conditional Generation Likelihood Advantage (CGLA). Our findings highlight the persistence of reliance on spurious correlations and the difficulty of mitigation on our benchmark. We hope this work can inspire new technical strides to mitigate these biases. Our benchmark is publicly available at https://huggingface.co/datasets/mmbench/MM-SpuBench.

本文貢獻

A benchmark for spurious biases in multimodal LLMs: paired image sets where the correct answer is independent of spurious cues (backgrounds, text overlays, co-occurrences). Reveals that strong MLLMs lean heavily on shortcuts, motivating bias-aware evaluation and mitigation.

BibTeX

@inproceedings{ye2026mmspubench,
  title     = {MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs},
  author    = {Wenqian Ye and Bohan Liu and Guangtao Zheng and Di Wang and Yunsheng Ma and Xu Cao and Bolin Lai and James M. Rehg and Aidong Zhang},
  booktitle = {ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2406.17126}
}