Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
Sanchit Sinha, Guangzhi Xiong, 刘博涵 (Bohan Liu), Zhenghao He, Aidong Zhang
摘要
此摘要由英文原文自动翻译。
思维链(CoT)提示在多模态大语言模型(MLLM)中的有效性仍不明确:在多个视觉推理基准上,CoT提示相较直接提示常导致性能下降。本文针对三个现代MLLM家族、跨模型规模,在需要逐步视觉证据的数据集上系统分析CoT行为。分析找出两个反复出现的失效模式:过早承诺答案,以及生成理由过程中对视觉token的直接访问受限。我们进一步发现,标准的CoT式有监督微调(CoT-SFT)仅能部分缓解这些问题,且常增加对文本先验的依赖、降低反事实视觉依赖。基于这些发现,我们提出Attentive-CoT(Att-CoT)——一个注意力引导的微调目标,促使CoT轨迹延迟答案承诺,同时保持对视觉token的持续访问。Att-CoT无需架构修改即可嵌入任何CoT-SFT训练。在三个视觉推理基准与六个MLLM上的实验表明,Att-CoT较标准微调提升了CoT性能。
原始摘要(英文)
The effectiveness of Chain-of-Thought (CoT) prompting in Multimodal Large Language Models (MLLMs) remains uncertain: across several visual reasoning benchmarks, CoT prompting often degrades performance compared to direct prompting. In this paper, we provide a systematic analysis of CoT behavior in three modern MLLM families across model scales on datasets requiring step-wise visual evidence. Our analysis identifies two recurring failure modes: premature answer commitment and limited direct visual-token access during rationale generation. We further find that standard CoT-style Supervised Fine-Tuning (CoT-SFT) can mitigate these issues only partially, while often increasing reliance on textual priors and reducing counterfactual visual dependence. Motivated by these findings, we propose Attentive-CoT (Att-CoT), an attention-guided fine-tuning objective that encourages CoT trajectories to delay answer commitment while maintaining sustained visual-token access. Att-CoT can be plugged into any CoT-SFT training run without architectural changes. Experiments on three visual reasoning benchmarks across six MLLMs show that Att-CoT enhances CoT performance over standard fine-tuning.
本文贡献
仅微调多模态大语言模型的少数注意力图(而非权重)即可改善思维链推理。注意力引导的更新在保留模型知识的同时控制其“关注哪里”,以远低于完全微调的成本提升ChartQA与CoT基准。
BibTeX
@misc{sinha2026attention,
title = {Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning},
author = {Sanchit Sinha and Guangzhi Xiong and Bohan Liu and Zhenghao He and Aidong Zhang},
year = {2026},
eprint = {arXiv:2606.01558},
howpublished = {arXiv:2606.01558}
}