Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Guangzhi Xiong, Zhenghao He, 刘博涵 (Bohan Liu), Sanchit Sinha, Aidong Zhang
摘要
此摘要由英文原文自动翻译。
检索增强生成(RAG)通过将输出锚定于检索证据来提升大语言模型(LLM)的事实性,但忠实性失败——生成内容与来源矛盾或超出来源——仍是关键挑战。现有的RAG幻觉检测方法多依赖需要大量标注数据的检测器大规模训练,或查询外部LLM裁判而导致高推理成本。虽有方法尝试利用LLM内部表征进行幻觉检测,其准确性仍有限。受机制可解释性最新进展启发,我们运用稀疏自编码器(SAE)解耦内部激活,成功找出在RAG幻觉期间被特定触发的特征。在基于信息量的特征选择与加法特征建模的系统化流程之上,我们提出RAGLens——一个轻量级幻觉检测器,利用LLM内部表征精确标记不忠实的RAG输出。RAGLens不仅检测性能优于现有方法,还为其判定提供可解释的理由,实现对不忠实RAG的有效事后缓解。最后,我们论证设计选择并揭示LLM内部幻觉相关信号分布的新见解。代码可在https://github.com/Teddy-XiongGZ/RAGLens获取。
原始摘要(英文)
Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generations contradict or extend beyond the provided sources, remain a critical challenge. Existing hallucination detection methods for RAG often rely either on large-scale detector training, which requires substantial annotated data, or on querying external LLM judges, which leads to high inference costs. Although some approaches attempt to leverage internal representations of LLMs for hallucination detection, their accuracy remains limited. Motivated by recent advances in mechanistic interpretability, we employ sparse autoencoders (SAEs) to disentangle internal activations, successfully identifying features that are specifically triggered during RAG hallucinations. Building on a systematic pipeline of information-based feature selection and additive feature modeling, we introduce RAGLens, a lightweight hallucination detector that accurately flags unfaithful RAG outputs using LLM internal representations. RAGLens not only achieves superior detection performance compared to existing methods, but also provides interpretable rationales for its decisions, enabling effective post-hoc mitigation of unfaithful RAG. Finally, we justify our design choices and reveal new insights into the distribution of hallucination-related signals within LLMs. The code is available at https://github.com/Teddy-XiongGZ/RAGLens.
本文贡献
RAG模型生成时常忽略检索到的证据。我们利用稀疏自编码器分解生成行为,找出影响忠实性的特征并加以引导——无需重新训练基础模型即可改善归因与依据。
BibTeX
@inproceedings{xiong2026faithful,
title = {Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders},
author = {Guangzhi Xiong and Zhenghao He and Bohan Liu and Sanchit Sinha and Aidong Zhang},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://arxiv.org/abs/2512.08892}
}