SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias

Wenqian Ye, Di Wang, Guangtao Zheng, 刘博涵 (Bohan Liu), Aidong Zhang

AAAI 2026 arXiv ↗ Scholar ↗

概要

この概要は英語原文から自動翻訳されています。

CLIPのような大規模視覚言語モデルは、画像とテキストを共有埋め込み空間で整列させることで、強力なゼロショット分類性能を示してきた。しかし、CLIPモデルはしばしばマルチモーダルなスプリアスバイアス、すなわち見かけの特徴に依存する望ましくない傾向を発達させる。例えば、CLIPは物体の本質的特徴ではなく、頻繁に共起する背景に基づいて画像中の物体タイプを推論することがある。このバイアスは、そのようなクロスモダルな関連が成り立たない分布外データにおいて、事前訓練されたCLIPモデルの頑健性を著しく損なう。既存のマルチモーダルスプリアスバイアス緩和手法は通常、下流データでのファインチューニングまたはバイアスに関する事前知識を必要とし、CLIPの箱から出してすぐの使える性質を損なう。本稿ではまず、ゼロショット分類におけるマルチモーダルスプリアスバイアスの影響を理論的に分析する。この知見に基づき、ガイド付きプロンプト選択によりスプリアスバイアスを緩和するシンプルかつ効果的な手法Spuriousness-Aware Guided Exploration(SAGE)を提案する。SAGEは訓練もファインチューニングも外部注釈も不要である。プロンプトテンプレートの空間を探索し、クラス間の意味的分離を最大にするプロンプトを選択することで、最悪グループ頑健性を向上させる。4つの実世界ベンチマークデータセットと5つの人気バックボーンモデルでの広範な実験により、SAGEは一貫してゼロショット性能と汎化を向上させ、外部知識やモデル更新なしに従来のゼロショット手法を上回ることを示す。

原文(英語)

Large vision-language models, such as CLIP, have shown strong zero-shot classification performance by aligning images and text in a shared embedding space. However, CLIP models often develop multimodal spurious biases, which is the undesirable tendency to rely on spurious features. For example, CLIP may infer object types in images based on frequently co-occurring backgrounds rather than the object's core features. This bias significantly impairs the robustness of pre-trained CLIP models on out-of-distribution data, where such cross-modal associations no longer hold. Existing methods for mitigating multimodal spurious bias typically require fine-tuning on downstream data or prior knowledge of the bias, which undermines the out-of-the-box usability of CLIP. In this paper, we first theoretically analyze the impact of multimodal spurious bias in zero-shot classification. Based on this insight, we propose Spuriousness-Aware Guided Exploration (SAGE), a simple and effective method that mitigates spurious bias through guided prompt selection. SAGE requires no training, fine-tuning, or external annotations. It explores a space of prompt templates and selects the prompts that induce the largest semantic separation between classes, thereby improving worst-group robustness. Extensive experiments on four real-world benchmark datasets and five popular backbone models demonstrate that SAGE consistently improves zero-shot performance and generalization, outperforming previous zero-shot approaches without any external knowledge or model updates.

本論文の貢献

プロンプトレベルのマルチモーダルバイアス防御:SAGEはプロンプトの「スプリアス意識」の程度を推定し、モデルがデータセットのショートカットではなく真の視覚内容に依存するような誘導付きプロンプト変形を探索します。

BibTeX

@inproceedings{ye2026sage,
  title     = {SAGE: Spuriousness-Aware Guided Prompt Exploration for Mitigating Multimodal Bias},
  author    = {Wenqian Ye and Di Wang and Guangtao Zheng and Bohan Liu and Aidong Zhang},
  booktitle = {AAAI Conference on Artificial Intelligence},
  year      = {2026},
  url       = {https://arxiv.org/abs/2511.13005}
}