500xCompressor (1) A-MEM (1) ACE (1) Activation Oracle (1) Activation Patching (2) Agentic Memory (1) AlignSAE (1) Angular Steering (1) attention map (2) Attention Sink (2) Bayesian Optimization (15) Benchmark (3) BRIDGE (1) Circuit Discovery (18) CircuitLasso (1) ClassifSAE (1) context-grounded faithfulness (1) Context Adaptation (1) Context Compression (1) Context Engineering (2) Contextual Faithfulness (1) Critical Component (7) Critical Neuron (1) Data-driven Circuit Discovery (1) Data Augmentation (2) data contamination (4) DCD (1) Defending LLM (8) Determinantal Point Process (4) Dictionary Learning (3) DoLa (2) DSPy (1) Dynamic Cheatsheet (1) EAP-IG (1) ERASER (1) FaithEval (1) Faithfulness of NLE (4) Feature Absorption (1) GEPA (1) Gradient-Based Activation Steering (1) HAGD (1) HbBoPs (1) HDBO (1) Head (3) Head-level steering (1) High-Dimensional BO (1) ICAE (1) In-Context Learning (9) Induction Head (1) Jailbreak Attack (4) knowledge editing (4) Language-Specific Neurons (2) Language Confusion (2) Latent Bayesian Optimization (2) LLM-as-a-judge (3) LLM-as-a-jury (1) LLM-DCP (1) LLM Quantization (3) Massive Activation (2) Matryoshka SAE (1) mechanistic interpretability (35) MERA (1) MIPRO (1) Multi-Objective Optimization (3) Neuron (15) NMF (1) Ontology Generation (2) Ontology Learning (6) Optuna (1) orthogonal feature directions (1) Perplexity Filtering (1) Personalized LLM (3) Prompt Compression (9) Prompt Optimization (1) Prompt Selection (2) RAG (2) Rationale (1) ReAct (1) Relevance Patching (1) RelP (1) retrieval head (2) SAE (13) Safety Knowledge Neuron (1) Semi-NMF (1) Simple Adaptive Attacks (1) Size-Fidelity Paradox (1) SNMF (1) Sparse Feature Coactivation (1) Spherical Steering (1) SR-NLE (2) steering LM (27) Streaming LLM (1) Style Modulation Heads (1) Super Weight (1) Survey (1) Test-Time Training (2) TPE (2) Transcoder (1) Tree-Structured Parzen Estimator (1) Weight Patching (1)

LLM 논문에 관한 블로그입니다.

  • ** DCR: Quantifying Data Contamination in LLMs Evaluation (EMNLP 2025)

    ** DCR: Quantifying Data Contamination in LLMs Evaluation (EMNLP 2025)

    https://www.dropbox.com/scl/fi/bk05fuy78jw67cehwmfh0/emnlp25_DCR_Framework_for_Reliable_LLM_Evaluation.pdf?rlkey=763fzmi7syaqw2g8kpu39sqoq&dl=0 이 논문은 LLM 평가에서의 Benchmark Data Contamination (BDC) 문제를 정량적으로 측정하고, 오염을 반영하여 성능을 보정하는 DCR (Data Contamination Risk) 프레임워크를 제안합니다. 핵심 메시지는 다음과 같습니다: LLM의 높은 benchmark 성능이 실제 일반화 능력이 아니라, 사전 학습 중 평가 데이터 노출(오염) 때문일 수 있다. 따라서 성능을 그대로 믿어서는 안 되며, 오염을 정량화하고 보정해야 한다. 1. 문제…

  • LLM-based Zero-shot Triple Extraction for Automated Ontology Generation from Software Engineering Standards (arXiv 2025)

    LLM-based Zero-shot Triple Extraction for Automated Ontology Generation from Software Engineering Standards (arXiv 2025)

    이 논문은 소프트웨어 공학 표준(SES) 문서로부터 LLM 기반 zero-shot triple extraction을 통해 **자동 온톨로지 생성(AOG)**을 수행하는 워크플로우를 제안합니다. 1. 문제 설정과 기여 문제 정의 핵심 기여 2. 전체 워크플로우 구조 (Figure 1, p.2) Workflow는 두 층으로 구성됩니다: 3. 핵심 알고리즘: Ontology Scaffold 생성 Algorithm 1 (p.3) 목표: G = (V, E) 단계별 설명 (1) Sentence-level…

  • Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field (Information Processing & Management, 2026)

    1. 연구 배경 및 문제 정의 왜 중요한가? 연구 주제 온톨로지(ontology of research topics)는 다음과 같은 핵심 인프라입니다: 그러나 기존 온톨로지는: 최근 LLM이 zero-shot 추론 능력을 보이면서, “LLM이 온톨로지의 핵심 관계 추론을 대신할 수 있는가?” 라는 질문이 제기됨. 2. 연구 목표 이 논문은 다음 문제를 다룹니다: 두 연구 주제 간 semantic relation을 LLM이 정확히 판별할…

  • Ontology Generation using Large Language Models (ArXiv 2025)

    Ontology Generation using Large Language Models (ArXiv 2025)

    https://www.dropbox.com/scl/fi/x7nfxsb3fbus5pv3fhpl9/arxiv25_LLM_Ontology_Engineering.pdf?rlkey=82887evkl4h1d1n8vi24pl1o7&dl=0 1. 연구 배경 및 문제 설정 문제의식 2. 연구 질문 (RQs) 논문은 세 가지 질문을 다룸: 3. 핵심 기여 (1) 두 가지 새로운 Prompting 기법 제안 (2) 다차원 평가 프레임워크 제안 (3) Benchmark Dataset 구축 4. 핵심 개념 정리 Ontology 정의 Modelled CQ Minimal Ontology Module Superfluous Element 5. 방법론 5.1 Ontology 생성 방식…

  • Methodological Exploration of Ontology Generation with a Dedicated Large Language Model (Electronics 2025)

    Methodological Exploration of Ontology Generation with a Dedicated Large Language Model (Electronics 2025)

    https://www.dropbox.com/scl/fi/yk61v11k250gpt73c553q/electronics25_Hybrid_LLM_Ontology_Automation_for_Autonomous_Vehicles.pdf?rlkey=hgf60n03yzf0mrd8nyffrzgmt&dl=0 다음 논문은 LLM을 활용한 온톨로지 개발 방법론을 제안하고, 이를 자율주행 차량의 Driver–Vehicle Interface(DVI) 도메인에 적용한 연구입니다. 1. 연구 목적과 문제의식 핵심 질문 동기 전통적인 온톨로지 개발은: LLM은: 이 가능하지만, 문제가 존재합니다. 따라서 이 논문은 Human-in-the-loop 기반 LLM 온톨로지 개발 프로세스를 제안합니다  . 2. 전체 방법론 구조 (7단계) 논문에서 제안하는 프로세스는 아래와 같습니다  : Phase…

  • * LLMs4OL: Large Language Models for Ontology Learning (ISWC 2023)

    * LLMs4OL: Large Language Models for Ontology Learning (ISWC 2023)

    https://www.dropbox.com/scl/fi/ilsqqrtimqvhy1fpsu1or/arxiv23_LLMs4OL_LLM_Driven_Ontology_Learning.pdf?rlkey=i41zsfz6t5utf2i9rfl9cy9uh&dl=0 다음 논문은 LLMs4OL: Large Language Models for Ontology Learning (ISWC 2023)이며, LLM을 Ontology Learning(OL)에 직접 적용한 최초의 체계적 실험 연구입니다. 1. 문제 설정: 왜 LLM으로 Ontology Learning인가? Ontology Learning (OL)이란? 텍스트로부터 다음을 자동으로 추출하여 구조화하는 작업입니다: 전통적 OL은: 문제점: 논문 핵심 가설 LLM의 emergent capability가 Ontology Learning에도 적용될 수 있는가? LLM은: → 그렇다면 ontology…

  • ** End-to-End Ontology Learning with Large Language Models (NeurIPS 2024)

    ** End-to-End Ontology Learning with Large Language Models (NeurIPS 2024)

    https://www.dropbox.com/scl/fi/iti1p1d7g37qzqu2a99g1/nips24_OLLM_End-to-End_Ontology_Learning.pdf?rlkey=qgwxrjz7syzldupesjy53i56h&dl=0 아래 논문은 **End-to-End Ontology Learning with Large Language Models (NeurIPS 2024)**로, LLM을 이용해 *온톨로지(특히 taxonomic backbone)*를 subtask 분해 없이 end-to-end로 학습하는 방법인 OLLM을 제안합니다  1. 문제 설정 Ontology Learning (OL)이란? 온톨로지는 예: 기존 OL 접근: 즉, subtask 조합 방식. 한계 2. 핵심 아이디어: OLLM 핵심 전환 “엣지를 예측하지 말고, subgraph 전체를 생성하자.” OLLM은 다음을…

  • ** Searching for Optimal Solutions with LLMs via Bayesian Optimization (ICLR 2025)

    ** Searching for Optimal Solutions with LLMs via Bayesian Optimization (ICLR 2025)

    https://www.dropbox.com/scl/fi/jcmsttjbcdo6e5fc7z2qt/iclr25_BOPRO_Bayesian_LLM_Search.pdf?rlkey=5sob9aj5n0hf3eplt4sohxr6f&dl=0 1. 문제의식: LLM 기반 “탐색”의 한계 최근 LLM을 테스트 타임에서 여러 번 샘플링하여 더 나은 해를 찾는 방식(test-time compute scaling)이 주목받고 있습니다. 하지만 기존 방식들은 다음 한계를 가집니다: 접근 한계 Repeated Sampling 탐색 공간 구조를 고려하지 않음 Greedy OPRO exploitation 위주 → local optima에 갇힘 진화 알고리즘 비용 큼 / 정적 전략 난이도 예측…

  • * Inversion-based Latent Bayesian Optimization (NeurIPS 2024)

    * Inversion-based Latent Bayesian Optimization (NeurIPS 2024)

    https://www.dropbox.com/scl/fi/tcikm6h4oilq8k599jbsr/nips24_InvBO_Latent_Bayesian_Optimization.pdf?rlkey=cxseh9vvsb81eu2i269b7mrt4&dl=0 논문의 핵심은 Latent Bayesian Optimization (LBO)의 misalignment 문제를 inversion으로 해결하고, trust region anchor 선택을 개선하는 것입니다  . 1. 문제 배경: Latent Bayesian Optimization (LBO) 1.1 LBO의 기본 구조 LBO는 이산/구조적 입력 공간(예: 분자, 수식)을 연속 latent space로 매핑한 뒤 BO를 수행합니다. Surrogate model g(z)는 사실상 f∘pθ:Z→Yf \circ p_\theta : Z \to Y 를 근사해야…

  • ** Hyperband-based Bayesian Optimization for Black-box Prompt Selection (ICML 2025)

    ** Hyperband-based Bayesian Optimization for Black-box Prompt Selection (ICML 2025)

    https://www.dropbox.com/scl/fi/hum1vhhz0g7patuqsz8wr/icml25-HbBoPs_Efficient_Prompt_Selection.pdf?rlkey=dmdudu6tl9z8bkeck6uk1xfap&dl=0 1. 문제 설정: Static Black-box Prompt Selection 목표 수식적으로는: arg⁡minp∈P⁡𝔼(x,y)[l(y,hp(x))]\arg\min_{p \in P} \mathbb{E}_{(x,y)}[l(y, h_p(x))] 하지만 실제로는 validation set 평균으로 근사: f(p)=1nvalid∑i=1nvalidl(yi,hp(xi))f(p) = \frac{1}{n_{valid}} \sum_{i=1}^{n_{valid}} l(y_i, h_p(x_i)) 여기서 핵심 제약은: 즉, 샘플 효율(sample-efficient) + 쿼리 효율(query-efficient) 이 동시에 필요함. 2. 기존 방법들의 한계 논문에서 지적한 문제점: 방법 한계 EASE exemplar selection 위주, 구조 정보 활용…

  • ** INSTRUCTZERO: Efficient Instruction Optimization for Black-Box Large Language Models (ICML 2024)

    ** INSTRUCTZERO: Efficient Instruction Optimization for Black-Box Large Language Models (ICML 2024)

    https://www.dropbox.com/scl/fi/2hdvce7ugxk8424cbf6su/icml24_InstructZero_Black_Box_Optimization.pdf?rlkey=5cpkzhft45p8xhpsak6c8czky&dl=0 1. 문제 정의: 왜 Instruction 최적화가 어려운가? LLM은 instruction-following 능력이 있지만, instruction phrasing에 매우 민감합니다. 동일한 의미라도 표현이 조금만 달라지면 성능이 크게 변합니다. 논문은 다음 문제를 다룹니다: maxv∈𝒱⁡𝔼(X,Y)∼Dth(f([v;X]),Y)\max_{v \in \mathcal{V}} \mathbb{E}_{(X,Y)\sim D_t} h(f([v;X]), Y) 핵심 난점 2. 핵심 아이디어 직접 instruction을 최적화하지 않는다. 대신, Soft prompt를 최적화해서, open-source LLM이 좋은 instruction을 생성하도록 유도한다. 전체…

  • * Bayesian Optimization for Instruction Generation (BOInG) (Applied Sciences, 2024)

    * Bayesian Optimization for Instruction Generation (BOInG) (Applied Sciences, 2024)

    다음 논문은 BO를 이용해 instruction(프롬프트)를 자동 생성하는 방법을 제안한 연구입니다: Sabbatella et al., “Bayesian Optimization for Instruction Generation (BOInG)”, Applied Sciences, 2024  1. 문제 설정: 왜 Instruction을 BO로 최적화하는가? LLM의 성능은 **instruction(=프롬프트)**에 매우 민감합니다. 특히 **black-box LLM (예: GPT-3.5, GPT-4o)**에서는 gradient 접근이 불가능하므로, instruction 최적화는 black-box combinatorial optimization 문제가 됩니다. 논문은 이를 다음과 같이 정식화합니다…

  • * Bayesian Optimization for Controlled Image Editing via LLMs (Findings of ACL 2025)

    * Bayesian Optimization for Controlled Image Editing via LLMs (Findings of ACL 2025)

    https://www.dropbox.com/scl/fi/vvpovz1rqzps555lwsjrs/facl25_BayesGenie_Autonomous_Image_Editing.pdf?rlkey=jusmqgj1e3b1hdhhezko9fk3r&dl=0 본 논문은 BayesGenie라는 프레임워크를 제안합니다. 핵심 아이디어는 다음과 같습니다: LLM을 “Promptist + Evaluator”로 사용하고, Bayesian Optimization(BO)을 통해 diffusion 모델의 CFG 파라미터를 자동 최적화하여 mask 없이 정밀한 이미지 편집을 수행한다. 1. 문제 설정 기존 한계 기존 image editing 방법들의 문제점: 2. BayesGenie 전체 구조 시스템 개요 (논문 Figure 2, p.4) 구조는 다음 4단계로 구성됩니다: ①…

  • ** (Latent) Bayesian Optimization(베이지안 최적화, BO)

    아래는 Bayesian Optimization(베이지안 최적화, BO) 를 직관 → 구성요소 → 알고리즘 흐름 → 예시 → 실전 팁/주의점 순서로 설명한 내용입니다. 1) BO가 풀고 싶은 문제는 뭔가? BO는 보통 이런 상황에서 씁니다. 즉, “적은 횟수의 실험으로 최적점을 찾는” 최적화. 2) 직관: “지도(확률모델) + 다음에 어디 볼지(탐색전략)” BO는 매 반복마다 딱 두 가지를 합니다. 이걸 반복합니다. 3)…

  • *** GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs (ArXiv 2024)

    *** GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs (ArXiv 2024)

    https://www.dropbox.com/scl/fi/ya2axdo5gka9awivkp87w/arxiv24_GASP_Stealthy_LLM_Jailbreaking.pdf?rlkey=qbleaiprrmcu5tx83w4osrtck&dl=0 아래는 GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs (arXiv 2024) 논문의 핵심 내용을 정리한 설명입니다  1. 문제 정의 및 동기 LLM은 RLHF 등으로 안전 정렬(alignment)이 되어 있지만, adversarial prompt (jailbreak) 를 통해 유해 응답을 유도할 수 있습니다. 기존 jailbreak 방법의 한계: 방법 한계 Heuristic (role-play 등) 일반화 어려움, 수작업 의존 GCG류…

  • * Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (NeurIPS 2024)

    * Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (NeurIPS 2024)

    https://www.dropbox.com/scl/fi/jd0im5p82nhdytkkpqfd8/nips24_TAP_Automated_LLM_Jailbreaking.pdf?rlkey=5eo4v3cihdxd3mthswnxywiv9&dl=0 논문 개요 이 논문은 black-box 환경에서 자동으로 LLM을 jailbreak하는 방법인 TAP (Tree of Attacks with Pruning) 을 제안합니다. 핵심은 다음 세 가지 조건을 모두 만족하는 공격 방법입니다: 기존 black-box 방법(PAIR)을 확장하여, branching + pruning 구조를 도입해 성공률을 크게 개선합니다  . 1. 문제 정의 LLM alignment (RLHF, guardrail 등)에도 불구하고, “How to build a bomb?”…

  • *** AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs (ICML 2025)

    *** AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs (ICML 2025)

    https://www.dropbox.com/scl/fi/kwj2led25zucha4tf3xoe/icml25_AdvPrompter_Fast_Adaptive_Adversarial_Prompting.pdf?rlkey=5opk654xmgml6cl1bje1xrpzw&dl=0 논문 **“AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs” (ICML 2025)**는 자동화된 adversarial red-teaming을 위한 LLM 기반 기법인 AdvPrompter를 제안합니다. 이 모델은 human-readable한 adversarial suffixes를 빠르게 생성하여 Target LLM을 jailbreak하는 데 사용됩니다. 아래는 논문의 핵심 내용입니다. 배경 및 문제의식 핵심 기여 1. AdvPrompter  모델 2. AdvPrompterTrain  (훈련 알고리즘) 3. AdvPrompterOpt  (suffix 생성 알고리즘) 실험 및 결과 ✔ 공격…

  • * Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing (IJCNLP-AACL 2025)

    * Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing (IJCNLP-AACL 2025)

    https://www.dropbox.com/scl/fi/81t9oact4pis3a5fsgsqw/ijcnlp25_SemanticSmooth_LLM_Defense.pdf?rlkey=v7wypnwghb1yeg7xlm5szdynx&dl=0 다음 논문은 IJCNLP-AACL 2025에 게재된 “Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing” 입니다  1. 문제 설정: Jailbreak 공격과 기존 방어의 한계 Jailbreak 공격이란? 정렬(aligned)된 LLM이 유해하거나 금지된 내용을 생성하도록 우회시키는 공격입니다. 논문에서는 다음과 같이 정의합니다: 공격 목표: JUDGE(F(x′))=1\text{JUDGE}(F(x’)) = 1 즉, 원래는 거부해야 할 유해 프롬프트를 수정해 수락하게 만드는 것 공격…

  • ** LLM Unlearning

    아래는 ACL/EMNLP/NAACL/COLING/NeurIPS/ICLR 학회의 unlearning 논문들의 방법론을 “비슷한 계열끼리 묶어서” 정리한 것입니다. 1) Forget/Retain 세트를 두고 “미세조정(FT)”로 지우는 계열 핵심 아이디어: 1-A. Gradient ascent / gradient difference 기반 이 계열은 구현이 단순하지만, (i) 과잉 삭제로 일반 성능 붕괴, (ii) 부분 삭제로 leakage, (iii) ‘지운 것 같은데 재학습/재노출에 취약’ 문제가 반복됩니다. 1-B. Retention을 “증류/보존”으로 강하게 잡는 계열 2) “Preference / Refusal…

  • ** Neuron-Level Knowledge Attribution in Large Language Models (EMNLP 2024)

    ** Neuron-Level Knowledge Attribution in Large Language Models (EMNLP 2024)

    https://www.dropbox.com/scl/fi/aqdb67mns9fzhmn3grtt9/emnlp24_LLM_-_.pdf?rlkey=zsjqywa861gc88il4s1tuvsdz&dl=0 아래는 EMNLP 2024 논문 “Neuron-Level Knowledge Attribution in Large Language Models” 의 핵심 내용을 정리한 설명입니다. 논문 개요 이 논문은 LLM 내부에서 특정 지식(facts)이 어떤 뉴런(neuron)에 저장되는지 정량적으로 찾아내는 뉴런 수준(neuron-level) attribution 방법을 제안합니다. 피쳐 단위(head, layer)보다 더 미세한 수준입니다. 기존 기법은 논문은 이를 해결하기 위해: 을 수행합니다. 배경 (왜 뉴런 수준인가?) 이전 연구들(Geva…

  • * Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering (ArXiv 2024)

    * Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering (ArXiv 2024)

    https://www.dropbox.com/scl/fi/iusmemx5gcdeakmjdx1mf/arxiv24_Peak_Activation_Ablation_Refined.pdf?rlkey=vec9sbmc0suns5j1fu1jy5gh6&dl=0 논문 “Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering” (Pochinkov et al., 2024) 은 Transformer 기반 모델의 주의(attention) 뉴런 해석과 절제(ablation) 방법을 체계적으로 비교하고, 새로운 방식인 **Peak Ablation (정점 중심 절제)**을 제안한 연구입니다. 아래에 핵심 내용을 구조적으로 정리했습니다. 1. 연구 배경 및 문제의식 기존 절제 방식: 2. 제안 개념: Peak…

  • *** Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis (ACL 2025)

    *** Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis (ACL 2025)

    https://www.dropbox.com/scl/fi/50rg8jb42dmsn3z0ssd46/acl25_Shortcut_Neurons_for_Trust.pdf?rlkey=2ob5vxseeql874rys9zxu7d53&dl=0 논문 “Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis” (ACL 2025) 은 데이터 오염(data contamination) 문제로 인해 LLM 평가의 신뢰성이 손상되는 문제를 해결하기 위해, 모델 내부의 “지름길 뉴런(shortcut neurons)”을 분석하고 억제함으로써 공정하고 신뢰할 수 있는 평가를 수행하는 방법을 제안한 연구입니다. 아래는 주요 내용 요약입니다. 연구 배경 및 문제의식 따라서 이 논문은 모델 내부의 원인,…

  • *** Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps (EMNLP 2024)

    *** Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps (EMNLP 2024)

    https://www.dropbox.com/scl/fi/pvl6q1znfir5lpw9qkcl9/emnlp24_Lookback_Lens_Tracking_LLM_Hallucinations.pdf?rlkey=gc8d7bbkyy18y4rftzwpih82d&dl=0 다음 논문은 “Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps” (EMNLP 2024) 입니다  . 이 논문은 **LLM의 contextual hallucination(문맥 기반 환각)**을 attention map만을 사용해 탐지하고, decoding 단계에서 이를 완화하는 방법을 제안합니다. 1. 문제 정의: Contextual Hallucination 논문은 환각을 두 종류로 구분합니다: 이 논문은 **후자(context-grounded setting)**에 집중합니다. 대표…

  • * Aging with GRACE: Lifelong Model Editing with Discrete Key-Value Adaptors (NeurIPS 2023)

    * Aging with GRACE: Lifelong Model Editing with Discrete Key-Value Adaptors (NeurIPS 2023)

    https://www.dropbox.com/scl/fi/5cjozna5e4uenu6rktf2t/nips23_Aging_With_GRACE.pdf?rlkey=ep237xt1wip67ep3si1hby5m8&dl=0 본 논문은 대규모 사전학습 모델을 재학습 없이, 수천 번 순차적으로(edit sequentially) 수정하는 방법을 제안합니다. 핵심은 모델 가중치를 건드리지 않고, 특정 레이어에 discrete key-value adaptor (codebook) 를 추가하여 “국소적 수정(spot-fix)”을 수행하는 것입니다  . 1. 문제 배경: 왜 Lifelong Model Editing이 필요한가? 배포된 LLM은 시간이 지나면서: 와 같은 문제가 발생합니다  . 그러나: → 따라서 문제는: 수백~수천…