[태그:] mechanistic interpretability
-

* A Mathematical Framework for Transformer Circuits (Transformer Circuits 2021)
https://www.dropbox.com/scl/fi/mmqtujkofh68ref3obbul/transformer_circuits21_Transformer_Circuit_Blueprint.pdf?rlkey=8hkvnezi1dhrmw3jlolcorrn4&dl=0 이 논문은 오늘날 Mechanistic Interpretability 분야의 출발점 중 하나로 평가받습니다. 특히 이후의 등의 연구들이 사실상 이 논문의 수학적 프레임워크 위에서 발전되었습니다. 1. 논문의 핵심 질문 Transformer 내부를 회로(circuit)처럼 해석할 수 있는가? 기존 Transformer 수식: Q=XWQQ=XW_Q K=XWKK=XW_K V=XWVV=XW_V A=softmax(QKT)A=\text{softmax}(QK^T) Y=AVWOY=AVW_O 은 학습과 구현에는 편하지만, “이 head가 실제로 무엇을 하는가?” 를 이해하기 어렵습니다. 저자들은 Transformer를…
-

* Circuit Breaking: Removing Model Behaviors with Targeted Ablation (ArXiv 2023)
https://www.dropbox.com/scl/fi/yjgk0u5i657adx5ub8633/arxiv23_Surgical_AI_Control.pdf?rlkey=39h35f7wx4hmekefdginb1hxx&dl=0 이 논문은 “모델의 특정 행동(behavior)만 제거할 수 있는가?” 라는 질문을 다룬다. 기존에는 Fine-tuning, RLHF, Model Editing 등이 주로 weight를 수정했는데, 이 논문은 훨씬 Mechanistic Interpretability 관점에서 접근한다. 핵심 아이디어는: “나쁜 행동을 만드는 circuit 전체를 찾는 대신, 그 circuit을 끊어버리는 최소 edge cut을 찾자.” 이다. 1. 문제 정의 논문은 “behavior removal”을 다음과 같이 정의한다.…
-

* IPE: Isolating Path Effects for Improving Latent Circuit Identification (BlackboxNLP 2025)
https://www.dropbox.com/scl/fi/bbgp8jxq3rqzki4hj11eg/blackboxnlp25_Isolating_Path_Effects.pdf?rlkey=9m7a2gec6dyfe4t5g2dqni5ee&dl=0 아래 논문은 IPE: Isolating Path Effects for Improving Latent Circuit Identification입니다. 핵심은 기존 circuit discovery가 edge 단위로 중요도를 계산하는 반면, 이 논문은 입력 임베딩 → 중간 컴포넌트들 → 최종 logits까지 이어지는 전체 computational path의 효과를 직접 분리해서 평가한다는 점입니다. 1. 문제의식 기존 방법들, 예를 들어 Activation Patching, Edge Activation Patching, ACDC, EAP는 보통…
-

* Circuit Component Reuse Across Tasks in Transformer Language Models (ICLR 2024)
https://www.dropbox.com/scl/fi/frg1w6t5gqipjmzmintg7/iclr24_Universal_Transformer_Circuit_Reuse.pdf?rlkey=bm2ytapv4bwgsq30q59tyl0lm&dl=0 논문: “Circuit Component Reuse Across Tasks in Transformer Language Models” (ICLR 2024) 1. 핵심 주장 이 논문은 Transformer LM 내부의 circuit component가 특정 task 전용이 아니라, 서로 다른 task에서도 재사용될 수 있다는 것을 보인다. 저자들은 두 task를 비교한다. Task 요구 행동 IOI: Indirect Object Identification 문장에서 indirect object 이름을 예측 Colored Objects 문맥에…
-

* Finding Neurons in a Haystack: Case Studies with Sparse Probing (ArXiv 2023)
https://www.dropbox.com/scl/fi/2xl6vrrw9zzoi0fcwi5wi/arxiv23_X-Raying_the_Black_Box.pdf?rlkey=3wnl2yjc5qoc3sxpebjvunfin&dl=0 이 논문은 **“LLM 내부에서 특정 개념(feature)이 몇 개의 뉴런에 의해 표현되는가?”**를 체계적으로 분석한 연구이다. 특히 기존 probing 연구를 확장하여 Sparse Probe를 사용함으로써 특정 feature와 관련된 뉴런을 매우 정밀하게 찾고, 이를 통해 monosemantic neuron, polysemantic neuron, superposition 현상을 실증적으로 분석한다. 1. 연구 배경 Mechanistic Interpretability 분야에서는 오래전부터 다음 질문이 존재했다. “특정 뉴런 하나가 하나의…
-

** Function Vectors in Large Language Models (ICLR 2024)
https://www.dropbox.com/scl/fi/hotaolcedy6yd2o5b48sc/iclr24_LLM_Function_Vector_Anatomy.pdf?rlkey=f4g9u2icoi4p79p8tiyu9xfon&dl=0 논문: Function Vectors in Large Language Models, ICLR 2024. 핵심은 ICL prompt가 유도한 “작업 함수”가 LLM 내부의 특정 attention head 출력들의 합으로 벡터화되어 있으며, 이 벡터를 다른 문맥에 삽입하면 모델이 해당 작업을 수행한다는 주장입니다. 1. 핵심 아이디어 논문은 LLM이 few-shot ICL을 할 때 단순히 예시를 복사하거나 표면 패턴을 따르는 것이 아니라, 예시들로부터 “입력→출력…
-
* Interpretability Analysis of Arithmetic In-Context Learning in Large Language Models (EMNLP 2025)
이 논문은 “LLM이 arithmetic ICL(In-Context Learning)을 할 때 실제로 무엇을 배우는가?” 를 mechanistic interpretability 관점에서 분석한 연구입니다. 특히 기존 연구가 주로 2-operand arithmetic (a+b) 를 분석한 반면, 본 논문은 3-operand arithmetic (a+b+c) 를 대상으로 합니다. 논문의 핵심 결론은 다음 한 문장으로 요약됩니다. LLM은 ICE(In-Context Example)의 산술적 정답을 배우기보다는 ICE의 패턴(format, structure) 을 학습하여 문제를…
-

* On Relation-Specific Neurons in Large Language Models (EMNLP 2025)
https://www.dropbox.com/scl/fi/7pm0homk42j0ndddk25qy/emnlp25_Mapping_LLM_Relation_Neurons.pdf?rlkey=6hz21la0g3x4mti2fk4q23pt8&dl=0 이 논문은 **“LLM 내부에 특정 사실(fact)을 저장하는 neuron이 아니라, 특정 관계(relation) 자체를 처리하는 neuron이 존재하는가?”**를 분석한 연구입니다. 기존 연구의 Knowledge Neuron은 (NVIDIA, CEO, Jensen Huang) 이라는 사실 전체를 저장하는 뉴런을 찾으려 했습니다. 반면 이 논문은 CEO 관계 자체를 담당하는 neuron 즉, 처럼 entity가 달라도 공통적으로 활성화되는 Relation-Specific Neuron (RelSpec Neuron) 이 존재하는지를 탐구합니다.…
-

*** Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps (EMNLP 2024)
https://www.dropbox.com/scl/fi/pvl6q1znfir5lpw9qkcl9/emnlp24_Lookback_Lens_Tracking_LLM_Hallucinations.pdf?rlkey=gc8d7bbkyy18y4rftzwpih82d&dl=0 다음 논문은 “Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps” (EMNLP 2024) 입니다 . 이 논문은 **LLM의 contextual hallucination(문맥 기반 환각)**을 attention map만을 사용해 탐지하고, decoding 단계에서 이를 완화하는 방법을 제안합니다. 1. 문제 정의: Contextual Hallucination 논문은 환각을 두 종류로 구분합니다: 이 논문은 **후자(context-grounded setting)**에 집중합니다. 대표…
-

** LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models (ArXiv 2025)
https://www.dropbox.com/scl/fi/85f1sbs7cy3jr2lhqi1l0/arxiv25_LogitLens4LLMs.pdf?rlkey=2rb3xyqhh3g22iwmnva1mi6wo&dl=0 아래는 **arXiv 2025 논문 *“LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models”***에 대한 설명입니다. 설명은 배경 → 방법론 → 시스템 설계 → 시각화 및 결과 → 기여와 한계 순으로 정리했습니다. 1. 연구 배경과 문제의식 Logit Lens는 중간 layer의 hidden state를 최종 LM head로 바로 투사하여, “이 layer에서 이미 어떤 토큰을 예측하고 있는가?”를 관찰하는 대표적인 mechanistic…