Towards Robustness against Typographic Attack with Training-free Concept Localization

刘博涵 (Bohan Liu), Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang

ECCV 2026 arXiv ↗ プロジェクト ↗ Scholar ↗

画像に描かれたテキストが、モデルの見ているものを決めてはいけません。注入された単語を読み取るCLIPの注意ヘッドの小さな集合を訓練なしで特定し、再重み付けまたは除去します。RTA-100(ViT-H/14)で物体認識精度+22.8ポイントを推論コストほぼゼロで達成し、5つのバックボーンで教師あり・訓練なし防御を上回ります。

概要

この概要は英語原文から自動翻訳されています。

Contrastive Language-Image Pretraining(CLIP)により訓練されたモデルは、現代のほとんどの大規模視覚言語モデル(LVLM)の基盤となる視覚エンコーダとして用いられている。広く採用されている一方で、CLIPモデルには重大かつ十分に検討されていない故障モードが存在する:画像内に現れる無関係なテキストが視覚表現を攪乱し、真の視覚意味論ではなく字句的意味へと偏らせる。この頑健性の問題はタイポグラフィック攻撃(Typographic Attack: TA)として知られ、自動運転などの安全が重要な応用に対する重大なリスクを露呈する。TAに対する解釈可能かつ効果的な頑健性を実現するため、我々は新規の訓練不要なメカニスティック解釈性手法を提案する。本手法は隠れ状態表現のサンプリングに基づく解釈を提供し、個々の注意ヘッドにおける意味的焦点と字句的焦点を定量的に帰属させる。確率論的解析とサーキットマイニングを通じて、字句情報を不釣り合いに多く符号化するVision Transformer(ViT)の特定コンポーネントを隔離し、TAのメカニスティックな発生源を特定する。さらに、追加訓練なしに特定されたサーキットへ直接適用する単純な介入が、物体分類におけるTA頑健性を大幅に向上させることを示す。注意重みの選択的調整などの介入は、教師ありおよび訓練不要な防御手法の両方を上回る。実験では、最先端の複数のLVLMの視覚エンコーダに提案する介入を適用すると、RIO-Bench上のTA干渉下でVisual Question Answering精度が大幅に向上することを示す。これらの結果は、我々のメカニスティックアプローチの有効性と汎化性の両方を裏付ける。コードはhttps://github.com/Liu-524/SamplingTARで公開されている。

原文(英語)

Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, commonly described as a Typographic Attack (TA), exposes a vulnerability that poses a significant risk to safety-critical applications such as autonomous driving. To achieve interpretable and effective robustness against TA, we propose a novel, training-free mechanistic interpretability method. Our method provides sampling-based interpretations of hidden state representations and quantitatively attributes semantic versus lexical focus to individual attention heads. Through probabilistic analysis and circuit mining, we isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, thereby identifying the mechanistic source of TA. We further show that simple interventions applied directly to the identified circuits, without any additional training, can substantially improve robustness against Typographic Attacks in object classification. These interventions, such as selective adjustment of attention weights, also outperform both supervised and training-free defense methods. Our experiments demonstrate that applying the proposed intervention to the vision encoders of several state-of-the-art LVLMs yields substantial gains in Visual Question Answering accuracy under Typographic Attack interference on RIO-Bench. These results confirm both the efficacy and the generalizability of our mechanistic approach. Code is released at https://github.com/Liu-524/SamplingTAR.

本論文の貢献

画像に描かれたテキストが、モデルの見ているものを決めてはいけません。注入された単語を読み取るCLIPの注意ヘッドの小さな集合を訓練なしで特定し、再重み付けまたは除去します。RTA-100(ViT-H/14)で物体認識精度+22.8ポイントを推論コストほぼゼロで達成し、5つのバックボーンで教師あり・訓練なし防御を上回ります。

BibTeX

@inproceedings{liu2026typographic,
  title     = {Towards Robustness against Typographic Attack with Training-free Concept Localization},
  author    = {Bohan Liu and Wenqian Ye and Guangzhi Xiong and Zhenghao He and Sanchit Sinha and Aidong Zhang},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2607.02494}
}