Towards Robustness against Typographic Attack with Training-free Concept Localization

刘博涵 (Bohan Liu), Wenqian Ye, Guangzhi Xiong, Zhenghao He, Sanchit Sinha, Aidong Zhang

ECCV 2026 arXiv ↗ 项目 ↗ Scholar ↗

画在图片上的文字不应决定模型看到什么。我们无需训练即可定位出读取注入文字的少量CLIP注意力头,并对其重新加权或消融。在RTA-100(ViT-H/14)上目标准确率提升22.8分,推理成本几乎为零,在五种骨干网络上超越有监督与免训练防御。

摘要

此摘要由英文原文自动翻译。

以对比语言-图像预训练(CLIP)训练的模型,是大多数现代大型视觉语言模型(LVLM)的基础视觉编码器。尽管被广泛采用,CLIP模型却存在一个关键且尚未充分探讨的失效模式:图像中出现的无关文字会混淆视觉表征,使其偏向词汇语义而非真正的视觉语义。这个鲁棒性问题一般称为打字攻击(Typographic Attack, TA),暴露出对自动驾驶等安全关键应用构成重大风险的漏洞。为了对TA实现可解释且有效的鲁棒性,我们提出一种新颖的免训练机制可解释性方法。本方法对隐藏状态表征提供基于采样的解释,并定量地将语义焦点与词汇焦点归因于各个注意力头。通过概率分析与电路挖掘,我们分离出不成比例地编码词汇信息的视觉Transformer(ViT)特定组件,从而找出TA的机制性根源。我们进一步证明,直接对所识别电路施加的简单干预——无需任何额外训练——即可大幅提升目标分类对打字攻击的鲁棒性。这些干预(例如选择性调整注意力权重)同时优于有监督与免训练防御方法。实验表明,将所提干预应用于多个最先进LVLM的视觉编码器,可在RIO-Bench的TA干扰下大幅提升视觉问答准确率。这些结果同时证实了我们机制性方法的有效性与泛化性。代码已在https://github.com/Liu-524/SamplingTAR公开。

原始摘要(英文)

Models trained via Contrastive Language-Image Pretraining (CLIP) serve as the foundational vision encoders for most modern Large Vision Language Models (LVLMs). Despite their widespread adoption, CLIP models exhibit a critical yet underexplored failure mode: irrelevant text appearing within images confounds visual representations, biasing them toward lexical meaning rather than true visual semantics. This robustness issue, commonly described as a Typographic Attack (TA), exposes a vulnerability that poses a significant risk to safety-critical applications such as autonomous driving. To achieve interpretable and effective robustness against TA, we propose a novel, training-free mechanistic interpretability method. Our method provides sampling-based interpretations of hidden state representations and quantitatively attributes semantic versus lexical focus to individual attention heads. Through probabilistic analysis and circuit mining, we isolate specific Vision Transformer (ViT) components that disproportionately encode lexical information, thereby identifying the mechanistic source of TA. We further show that simple interventions applied directly to the identified circuits, without any additional training, can substantially improve robustness against Typographic Attacks in object classification. These interventions, such as selective adjustment of attention weights, also outperform both supervised and training-free defense methods. Our experiments demonstrate that applying the proposed intervention to the vision encoders of several state-of-the-art LVLMs yields substantial gains in Visual Question Answering accuracy under Typographic Attack interference on RIO-Bench. These results confirm both the efficacy and the generalizability of our mechanistic approach. Code is released at https://github.com/Liu-524/SamplingTAR.

本文贡献

画在图片上的文字不应决定模型看到什么。我们无需训练即可定位出读取注入文字的少量CLIP注意力头,并对其重新加权或消融。在RTA-100(ViT-H/14)上目标准确率提升22.8分,推理成本几乎为零,在五种骨干网络上超越有监督与免训练防御。

BibTeX

@inproceedings{liu2026typographic,
  title     = {Towards Robustness against Typographic Attack with Training-free Concept Localization},
  author    = {Bohan Liu and Wenqian Ye and Guangzhi Xiong and Zhenghao He and Sanchit Sinha and Aidong Zhang},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2607.02494}
}