Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models
Zhenghao He, Guangzhi Xiong, 刘博涵 (Bohan Liu), Sanchit Sinha, Aidong Zhang
摘要
此摘要由英文原文自动翻译。
思维链(CoT)提示常能提升大语言模型(LLM)的推理性能,但触发该行为的内部信号仍未被充分理解。利用稀疏自编码器(SAE)捕捉的稀疏特征,我们提出一个系统化框架来分析并干预LLM的内部表征,识别一小簇与推理行为相关、且可通过定向干预进行因果检验的潜在特征。跨多个模型家族与推理基准,我们表明仅引导一个或少数推理相关潜在特征,即可在无显式CoT提示的情况下实质性诱发推理行为,达到与CoT相当的准确率。我们进一步表明,所识别的特征并不依赖特定措辞模式或冗长度,并通过抑制实验——即使在CoT提示下也损害性能——确认其在推理中的因果角色。这些结果表明,CoT提示通过激活特定潜在特征来触发推理,而对这些特征的定向干预提供了无需显式CoT即可引出高效推理行为的替代途径。代码可在https://github.com/Zhenghao-He/LatentCoT获取。
原始摘要(英文)
Chain-of-Thought (CoT) prompting often improves the reasoning performance of large language models (LLMs), but the internal signal that triggers this behavior remains poorly understood. Leveraging the sparse features captured by Sparse Autoencoders (SAEs), we propose a systematic framework to analyze and intervene on the internal representations of LLMs, identifying a small set of latent features that are linked to reasoning behavior and can be causally tested through targeted intervention. Across multiple model families and reasoning benchmarks, we show that steering one or a small number of reasoning-related latent features can substantially induce reasoning behavior without explicit CoT prompting, achieving accuracy comparable to CoT. We further show that the identified features are not tied to particular wording patterns or verbosity, and confirm their causal role in reasoning through suppression experiments that impair performance even under CoT prompting. These results suggest that CoT prompting activates specific latent features to trigger reasoning, and that targeted intervention on these features offers an alternative pathway to elicit efficient reasoning behavior without explicit CoT prompting. Code is available at https://github.com/Zhenghao-He/LatentCoT.
本文贡献
思维链并非大语言模型唯一的推理方式。我们发现一种潜在计算模式——推理在生成任何token之前即于隐藏状态中完成——并刻画潜在计算优于显式CoT的条件。
BibTeX
@misc{he2026latent,
title = {Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models},
author = {Zhenghao He and Guangzhi Xiong and Bohan Liu and Sanchit Sinha and Aidong Zhang},
year = {2026},
eprint = {arXiv:2601.08058},
howpublished = {arXiv:2601.08058}
}