[태그:] Self-Critique + Refinement

  • * Self-Critique and Refinement for Faithful Natural Language Explanations (EMNLP 2025)

    * Self-Critique and Refinement for Faithful Natural Language Explanations (EMNLP 2025)

    https://www.dropbox.com/scl/fi/ei33lu8j3vn1nmxuqbvnb/emnlp25_Faithful_AI_Explanations_via_SR-NLE.pdf?rlkey=o6pk8c120i861aiiz78866whb&dl=0 다음 논문은 LLM이 생성하는 자연어 설명(NLE)의 “faithfulness(충실성)”을 어떻게 개선할 것인가를 다룬 매우 중요한 연구입니다. 핵심은 모델이 스스로 자신의 설명을 비판하고 수정할 수 있는가입니다. 1. 문제 정의 (Why this paper?) 핵심 문제: NLE의 “비충실성 (Unfaithfulness)” 예: 즉, plausible explanation ≠ faithful explanation 2. 핵심 아이디어: SR-NLE Self-Critique + Refinement 논문에서 제안한 프레임워크: SR-NLE (Self-critique and…