Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models

Zhenghao He, Guangzhi Xiong, 刘博涵 (Bohan Liu), Sanchit Sinha, Aidong Zhang

arXiv 2026 arXiv ↗ Scholar ↗

Chain-of-thought is not the only way LLMs reason. We uncover a latent computational mode, reasoning that completes inside hidden states before any tokens are emit...

摘要

此摘要由英文原文自動翻譯。

思維鏈(CoT)提示常能提升大型語言模型(LLM)的推理效能,但觸發此行為的內部訊號仍未被充分理解。利用稀疏自編碼器(SAE)捕捉的稀疏特徵,我們提出一個系統化框架來分析並介入LLM的內部表徵,識別一小組與推理行為相關、且可透過標定介入進行因果檢驗的潛在特徵。跨多個模型家族與推理基準,我們顯示僅引導一個或少數推理相關潛在特徵,即可在無明示CoT提示的情況下實質誘發推理行為,達到與CoT相當的準確率。我們進一步顯示,所識別的特徵並不依賴特定措辭模式或冗長度,並透過抑制實驗——即使在CoT提示下也損害效能——確認其在推理中的因果角色。這些結果表明,CoT提示透過活化特定潛在特徵來觸發推理,而對這些特徵的標定介入提供了無需明示CoT即可引出高效推理行為的替代途徑。程式碼可於https://github.com/Zhenghao-He/LatentCoT取得。

原始摘要(英文)

Chain-of-Thought (CoT) prompting often improves the reasoning performance of large language models (LLMs), but the internal signal that triggers this behavior remains poorly understood. Leveraging the sparse features captured by Sparse Autoencoders (SAEs), we propose a systematic framework to analyze and intervene on the internal representations of LLMs, identifying a small set of latent features that are linked to reasoning behavior and can be causally tested through targeted intervention. Across multiple model families and reasoning benchmarks, we show that steering one or a small number of reasoning-related latent features can substantially induce reasoning behavior without explicit CoT prompting, achieving accuracy comparable to CoT. We further show that the identified features are not tied to particular wording patterns or verbosity, and confirm their causal role in reasoning through suppression experiments that impair performance even under CoT prompting. These results suggest that CoT prompting activates specific latent features to trigger reasoning, and that targeted intervention on these features offers an alternative pathway to elicit efficient reasoning behavior without explicit CoT prompting. Code is available at https://github.com/Zhenghao-He/LatentCoT.

本文貢獻

Chain-of-thought is not the only way LLMs reason. We uncover a latent computational mode, reasoning that completes inside hidden states before any tokens are emitted, and characterize when latent computation beats explicit CoT.

BibTeX

@misc{he2026latent,
  title         = {Triggering Chain-of-Thought via Latent Feature Interventions in Large Language Models},
  author        = {Zhenghao He and Guangzhi Xiong and Bohan Liu and Sanchit Sinha and Aidong Zhang},
  year          = {2026},
  eprint        = {arXiv:2601.08058},
  howpublished  = {arXiv:2601.08058}
}