Towards Robustness against Typographic Attack with Training-free Concept Localization

刘博涵 (Bohan Liu), Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang

ECCV 2026 arXiv ↗ 專案 ↗ Scholar ↗

Text painted on an image should not decide what the model sees. We localize, training-free, the small set of CLIP attention heads that read injected words, then r...

摘要

此摘要由英文原文自動翻譯。

以對比語言-影像預訓練(CLIP)訓練的模型,是大多數現代大型視覺語言模型(LVLM)的基礎視覺編碼器。儘管被廣泛採用,CLIP模型卻存在一個關鍵且尚未充分探討的失效模式:影像中出現的無關文字會混淆視覺表徵,使其偏向詞彙語意而非真正的視覺語意。這個強健性問題一般稱為打字攻擊(Typographic Attack, TA),暴露出對自駕等安全關鍵應用構成重大風險的漏洞。為了對TA實現可解釋且有效的強健性,我們提出一種新穎的免訓練機制可解釋性方法。本方法對隱藏狀態表徵提供基於抽樣的解釋,並定量地將語意焦點與詞彙焦點歸因於個別注意力頭。透過機率分析與電路挖掘,我們隔離出不成比例地編碼詞彙資訊的視覺Transformer(ViT)特定元件,從而找出TA的機制性根源。我們進一步證明,直接對所識別電路施加的簡單介入——無需任何額外訓練——即可大幅提升物體分類對打字攻擊的強健性。這些介入(例如選擇性調整注意力權重)同時優於監督式與免訓練防禦方法。實驗顯示,將所提介入應用於多個最先端LVLM的視覺編碼器,可在RIO-Bench的TA干擾下大幅提升視覺問答準確率。這些結果同時證實了我們機制性方法的有效性與泛化性。程式碼已公開於https://github.com/Liu-524/SamplingTAR。

原始摘要(英文)

Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, commonly described as a Typographic Attack (TA), exposes a vulnerability that poses a significant risk to safety-critical applications such as autonomous driving. To achieve interpretable and effective robustness against TA, we propose a novel, training-free mechanistic interpretability method. Our method provides sampling-based interpretations of hidden state representations and quantitatively attributes semantic versus lexical focus to individual attention heads. Through probabilistic analysis and circuit mining, we isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, thereby identifying the mechanistic source of TA. We further show that simple interventions applied directly to the identified circuits, without any additional training, can substantially improve robustness against Typographic Attacks in object classification. These interventions, such as selective adjustment of attention weights, also outperform both supervised and training-free defense methods. Our experiments demonstrate that applying the proposed intervention to the vision encoders of several state-of-the-art LVLMs yields substantial gains in Visual Question Answering accuracy under Typographic Attack interference on RIO-Bench. These results confirm both the efficacy and the generalizability of our mechanistic approach. Code is released at https://github.com/Liu-524/SamplingTAR.

本文貢獻

Text painted on an image should not decide what the model sees. We localize, training-free, the small set of CLIP attention heads that read injected words, then reweight or ablate them. +22.8 points object accuracy on RTA-100 (ViT-H/14) with near-zero test-time cost, beating supervised and training-free defenses across five backbones.

BibTeX

@inproceedings{liu2026typographic,
  title     = {Towards Robustness against Typographic Attack with Training-free Concept Localization},
  author    = {Bohan Liu and Wenqian Ye and Guangzhi Xiong and Zhenghao He and Sanchit Sinha and Aidong Zhang},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2607.02494}
}