Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Guangzhi Xiong, Zhenghao He, 刘博涵 (Bohan Liu), Sanchit Sinha, Aidong Zhang
概要
この概要は英語原文から自動翻訳されています。
検索拡張生成(RAG)は、検索された証拠に出力を接地させることで大規模言語モデル(LLM)の事実性を向上させるが、生成が提供された情報源と矛盾したり超出したりする忠実性の失敗は、依然として重大な課題である。既存のRAGのハルシネーション検出手法は、大規模なアノテーションデータを必要とする検出器の訓練か、推論コストの高い外部LLM判定者への照会のいずれかに依存しがちである。内部表現を利用する手法もあるが、その精度は限られている。メカニスティック解釈性の最新の進展に動機づけられ、我々はスパースオートエンコーダ(SAE)を用いて内部活性化を解きほぐし、RAGハルシネーション中に特異的に発火する特徴の特定に成功した。情報量に基づく特徴選択と加法的特徴モデリングの体系的パイプラインの上に、LLM内部表現を用いて不忠実なRAG出力を正確に検出する軽量なハルシネーション検出器RAGLensを導入する。RAGLensは既存手法を上回る検出性能に加え、判定根拠の解釈可能な説明を提供し、不忠実なRAGの効果的な事後緩和を可能にする。最後に、設計選択を正当化し、LLM内部におけるハルシネーション関連信号の分布に関する新たな知見を明らかにする。コードはhttps://github.com/Teddy-XiongGZ/RAGLensで入手可能である。
原文(英語)
Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generations contradict or extend beyond the provided sources, remain a critical challenge. Existing hallucination detection methods for RAG often rely either on large-scale detector training, which requires substantial annotated data, or on querying external LLM judges, which leads to high inference costs. Although some approaches attempt to leverage internal representations of LLMs for hallucination detection, their accuracy remains limited. Motivated by recent advances in mechanistic interpretability, we employ sparse autoencoders (SAEs) to disentangle internal activations, successfully identifying features that are specifically triggered during RAG hallucinations. Building on a systematic pipeline of information-based feature selection and additive feature modeling, we introduce RAGLens, a lightweight hallucination detector that accurately flags unfaithful RAG outputs using LLM internal representations. RAGLens not only achieves superior detection performance compared to existing methods, but also provides interpretable rationales for its decisions, enabling effective post-hoc mitigation of unfaithful RAG. Finally, we justify our design choices and reveal new insights into the distribution of hallucination-related signals within LLMs. The code is available at https://github.com/Teddy-XiongGZ/RAGLens.
本論文の貢献
RAGモデルは生成時に検索された証拠を無視しがちです。スパースオートエンコーダで生成挙動を分解し、忠実性を担う特徴を特定・操作することで、ベースモデルの再訓練なしに帰属と根拠付けを改善します。
BibTeX
@inproceedings{xiong2026faithful,
title = {Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders},
author = {Guangzhi Xiong and Zhenghao He and Bohan Liu and Sanchit Sinha and Aidong Zhang},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://arxiv.org/abs/2512.08892}
}