Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
Guangzhi Xiong, Zhenghao He, 刘博涵 (Bohan Liu), Sanchit Sinha, Aidong Zhang
초록
이 초록은 영어 원문에서 자동 번역되었습니다.
검색 증강 생성(RAG)은 검색된 근거에 출력을 기반함으로써 대규모 언어 모델(LLM)의 사실성을 높이지만, 생성물이 제공된 출처와 모순되거나 이를 넘어서는 충실성 실패는 여전히 중대한 과제다. 기존 RAG 환각 검출 방법은 대규모 주석 데이터가 필요한 검출기 학습이나 높은 추론 비용을 유발하는 외부 LLM 심판 조회에 의존하는 경우가 많다. LLM 내부 표현을 활용하려는 접근도 있지만 정확도가 제한적이다. 메커니즘 해석가능성의 최근 진전에 동기받아, 우리는 희소 오토인코더(SAE)를 활용해 내부 활성을 분해하고 RAG 환각 중 특이적으로 활성화되는 특징을 성공적으로 식별했다. 정보량 기반 특징 선택과 가법적 특징 모델링의 체계적 파이프라인 위에, LLM 내부 표현을 사용해 불충실한 RAG 출력을 정확히 플래그하는 경량 환각 검출기 RAGLens를 제안한다. RAGLens는 기존 방법을 능가하는 검출 성능과 함께 판정 근거의 해석 가능한 설명을 제공하여 불충실한 RAG의 효과적인 사후 완화를 가능하게 한다. 마지막으로 설계 선택을 정당화하고 LLM 내부의 환각 관련 신호 분포에 대한 새로운 통찰을 제시한다. 코드는 https://github.com/Teddy-XiongGZ/RAGLens에서 이용 가능하다.
원문 초록 (영어)
Retrieval-Augmented Generation (RAG) improves the factuality of large language models (LLMs) by grounding outputs in retrieved evidence, but faithfulness failures, where generations contradict or extend beyond the provided sources, remain a critical challenge. Existing hallucination detection methods for RAG often rely either on large-scale detector training, which requires substantial annotated data, or on querying external LLM judges, which leads to high inference costs. Although some approaches attempt to leverage internal representations of LLMs for hallucination detection, their accuracy remains limited. Motivated by recent advances in mechanistic interpretability, we employ sparse autoencoders (SAEs) to disentangle internal activations, successfully identifying features that are specifically triggered during RAG hallucinations. Building on a systematic pipeline of information-based feature selection and additive feature modeling, we introduce RAGLens, a lightweight hallucination detector that accurately flags unfaithful RAG outputs using LLM internal representations. RAGLens not only achieves superior detection performance compared to existing methods, but also provides interpretable rationales for its decisions, enabling effective post-hoc mitigation of unfaithful RAG. Finally, we justify our design choices and reveal new insights into the distribution of hallucination-related signals within LLMs. The code is available at https://github.com/Teddy-XiongGZ/RAGLens.
이 논문의 기여
RAG 모델은 생성 시 검색된 근거를 종종 무시합니다. 희소 오토인코더로 생성 행동을 분해하고, 충실성을 담당하는 특징을 찾아 조종합니다 — 기반 모델 재학습 없이 귀속과 근거성을 개선합니다.
BibTeX
@inproceedings{xiong2026faithful,
title = {Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders},
author = {Guangzhi Xiong and Zhenghao He and Bohan Liu and Sanchit Sinha and Aidong Zhang},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://arxiv.org/abs/2512.08892}
}