Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders

Guangzhi Xiong, Zhenghao He, 刘博涵 (Bohan Liu), Sanchit Sinha, Aidong Zhang

ICLR 2026 arXiv ↗ Scholar ↗

摘要

此摘要由英文原文自動翻譯。

檢索增強生成(RAG)透過將輸出錨定於檢索證據來提升大型語言模型(LLM)的事實性,但忠實性失敗——生成內容與來源矛盾或超出來源——仍是關鍵挑戰。現有的RAG幻覺偵測方法多依賴需要大量標註資料的檢測器大規模訓練,或查詢外部LLM裁判而導致高推論成本。雖有方法嘗試利用LLM內部表徵進行幻覺偵測,其準確性仍有限。受機制可解釋性最新進展啟發,我們運用稀疏自編碼器(SAE)拆解內部活化,成功找出在RAG幻覺期間被特定觸發的特徵。在基於資訊量的特徵選擇與加法特徵建模的系統化流程之上,我們提出RAGLens——一個輕量級幻覺偵測器,利用LLM內部表徵精確標記不忠實的RAG輸出。RAGLens不僅偵測效能優於現有方法,還為其判定提供可解釋的理由,實現對不忠實RAG的有效事後緩解。最後,我們論證設計選擇並揭示LLM內部幻覺相關訊號分布的新見解。程式碼可於https://github.com/Teddy-XiongGZ/RAGLens取得。

原始摘要(英文)

Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generations contradict or extend beyond the provided sources, remain a critical challenge. Existing hallucination detection methods for RAG often rely either on large-scale detector training, which requires substantial annotated data, or on querying external LLM judges, which leads to high inference costs. Although some approaches attempt to leverage internal representations of LLMs for hallucination detection, their accuracy remains limited. Motivated by recent advances in mechanistic interpretability, we employ sparse autoencoders (SAEs) to disentangle internal activations, successfully identifying features that are specifically triggered during RAG hallucinations. Building on a systematic pipeline of information-based feature selection and additive feature modeling, we introduce RAGLens, a lightweight hallucination detector that accurately flags unfaithful RAG outputs using LLM internal representations. RAGLens not only achieves superior detection performance compared to existing methods, but also provides interpretable rationales for its decisions, enabling effective post-hoc mitigation of unfaithful RAG. Finally, we justify our design choices and reveal new insights into the distribution of hallucination-related signals within LLMs. The code is available at https://github.com/Teddy-XiongGZ/RAGLens.

本文貢獻

RAG models often ignore retrieved evidence when generating. We use sparse autoencoders to decompose generation behavior, identify features responsible for faithfulness, and steer them, improving attribution and grounding without retraining the base model.

BibTeX

@inproceedings{xiong2026faithful,
  title     = {Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders},
  author    = {Guangzhi Xiong and Zhenghao He and Bohan Liu and Sanchit Sinha and Aidong Zhang},
  booktitle = {International Conference on Learning Representations (ICLR)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2512.08892}
}