Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains (ArXiv 2025)

https://www.dropbox.com/scl/fi/evnpykx30w6y77z6rymsg/arxiv25_Rubrics_as_Rewards.pdf?rlkey=bzu5tg8cgt0gk572zzwnv45lk&st=inpbgyxg&dl=0

이 논문의 핵심 아이디어는 매우 명확합니다.

수학·코딩처럼 정답을 자동 검증할 수 없는 영역에서도, 문제별(instance-specific) rubric을 여러 개의 작은 verifier처럼 사용하면 RLVR의 장점을 유지하면서 RL을 수행할 수 있다.

기존 RLVR(Reinforcement Learning with Verifiable Rewards)는 exact match, unit test 등 명확한 보상이 존재하는 수학·코딩에서 강력하지만, 의료 상담처럼 정확성·완전성·안전성·문맥 적합성 등이 동시에 필요한 open-ended 문제에서는 단일 binary reward를 만들기 어렵습니다. 저자들은 이를 **Rubrics as Rewards(RaR)**라는 구조화된 reward로 바꾸고 GRPO를 수행합니다. 가장 좋은 RaR-Implicit는 Direct-Likert reward 대비 HealthBench에서 최대 31% relative improvement, GPQA-Diamond에서 7% relative improvement를 보입니다.

아래에서는 관련 연구 → 방법론 → 실험 설계 → 실험 결과 → ablation → 논문의 의미와 한계 순으로 자세히 설명하겠습니다.


1. 문제의 출발점: 왜 Rubric이 필요한가?

1.1 기존 RLVR

RLVR의 전형적인 reward는 r(x,y^)={1,y^가 정답0,otherwiser(x,\hat y)= \begin{cases} 1,& \hat y\text{가 정답}\\ 0,& \text{otherwise} \end{cases}

처럼 정의할 수 있습니다.

예:

  • GSM8K: 숫자 final answer가 일치하는가?
  • MATH: symbolic/numeric answer가 정답인가?
  • Code: test case를 통과하는가?

따라서 별도의 learned reward model 없이도 ground-truth verifier로 RL을 할 수 있습니다.

하지만 의료 질문을 생각하면,

“이 환자에게 어떤 처치를 권해야 하는가?”

라는 답변에는 단순히 진단명 하나만 맞는 것으로 충분하지 않습니다.

예를 들어:

  1. 진단이 맞는가?
  2. 핵심 증거를 설명했는가?
  3. 위험한 처치를 권하지 않았는가?
  4. 중요한 differential diagnosis를 고려했는가?
  5. 설명이 충분한가?

를 모두 봐야 합니다.

논문은 RLVR의 이런 한계를 multi-criteria reward로 확장합니다. 기존 preference-based reward model은 response length나 formatting 같은 superficial artifact, annotator bias에 과적합할 수 있으며 대량의 pairwise preference data도 필요하다는 문제를 지적합니다.


2. 핵심 직관: Rubric = 여러 개의 작은 verifier

RaR의 가장 중요한 개념은 다음입니다.

기존 RLVR: Prompt→Response→single verifier→r\text{Prompt} \rightarrow \text{Response} \rightarrow \boxed{\text{single verifier}} \rightarrow r

RaR: Prompt→Response→{c1:diagnosis correct?c2:key sign explained?c3:safety issue covered?c4:common pitfall avoided?⋮→r\text{Prompt}\rightarrow\text{Response}\rightarrow\begin{cases}c_1 :\text{diagnosis correct?}\\c_2:\text{key sign explained?}\\c_3:\text{safety issue covered?}\\c_4:\text{common pitfall avoided?}\\\vdots\end{cases}\rightarrow r

즉 하나의 복잡한 “좋은 답인가?” 판단을 여러 개의 명시적인 subgoal로 분해합니다.

논문의 Figure 1도 정확히 이 구조입니다.

Reference answer→LLM rubric generation\boxed{\text{Reference answer}} \rightarrow \boxed{\text{LLM rubric generation}}

그 다음 x→πθ→y^1,…,y^16→LLM Judge + Rubric→R1,…,R16→GRPOx \rightarrow \pi_\theta \rightarrow \hat y_1,\ldots,\hat y_{16} \rightarrow \text{LLM Judge + Rubric} \rightarrow R_1,\ldots,R_{16} \rightarrow \text{GRPO}

입니다. Reference answer를 expert supervision의 proxy로 사용해 prompt-specific rubric을 생성하고, 그 rubric을 judge에게 주어 reward를 계산한 뒤 GRPO policy update를 수행합니다.


3. Related Work

논문의 Related Work는 크게 세 갈래입니다.


3.1 RLVR를 math/code 밖으로 확장하려는 연구

초기의 RLVR는 주로:

  • mathematical reasoning
  • code generation

에 집중했습니다.

논문에서 언급하는 최근 흐름은 다음과 같습니다.

General-Reasoner

GENERAL-REASONER는 약 200K개의 mixed-domain corpus를 이용하여:

  • physics
  • finance
  • policy

등으로 GRPO 기반 RLVR를 확장하고 MMLU-Pro에서 약 10점 향상을 보고합니다.

Cross-domain RLVR

후속 연구에서는:

  • medicine
  • chemistry
  • psychology
  • economics

까지 확장하고 하나의 cross-domain reward model이 여러 도메인을 supervise할 가능성을 보여줍니다.

Med-RLVR

MED-RLVR는 의료 multiple-choice QA에 verifiable reward를 적용하여 3B 모델에서도 reasoning capability를 끌어냅니다. 하지만 결국 MCQ이므로 답이 맞는지 자동 판별할 수 있다는 점에서는 여전히 RLVR-friendly task입니다.

RaR가 이들과 다른 지점은:

“verifiable domain을 넓힌다”가 아니라 “verifiability 자체가 약한 문제를 structured criteria로 바꾼다”는 것입니다.


3.2 Rubric을 이용한 LLM evaluation

두 번째 계열이 RaR와 가장 직접적으로 관련됩니다.

이미 rubric 기반 평가는 상당히 널리 연구되었습니다.

대표적으로:

HealthBench

의료 답변에 대해 clinician-written criteria를 사용합니다.

한 response에 대해 단순히

score = 7/10

으로 평가하지 않고,

  • communication quality
  • instruction following
  • accuracy
  • context awareness
  • completeness

등을 세부적으로 평가합니다.

HealthBench는 약 48K clinician-written criteria와 LLM judge를 결합합니다.

Rubric is All You Need

Pathak et al.은 question-specific rubric이 generic checklist보다 code evaluation의 accuracy와 consistency를 향상시킨다고 보고합니다.

LLM-Rubric

multidimensional calibrated evaluation이라는 방향입니다.


RaR와의 차이

기존 연구:Rubric→Evaluation\boxed{\text{Rubric}} \rightarrow \text{Evaluation}

RaR: Rubric→RL reward→policy learning\boxed{\text{Rubric}} \rightarrow \boxed{\text{RL reward}} \rightarrow \boxed{\text{policy learning}}

즉 rubric을 평가 도구에서 training-time reward function으로 승격시킨 것이 핵심 novelty입니다.

논문은 기존 rubric 연구가 주로 output grading 또는 DPO용 preference data conditioning에 사용된 반면, RaR는 이를 on-policy RL의 직접 reward로 사용한다고 강조합니다.


3.3 Preference Learning / RLHF / Process Reward

기존 RLHF는 ya≻yby_a \succ y_b

같은 pairwise preference를 학습하여 reward model Rϕ(x,y)R_\phi(x,y)을 구축합니다.

문제는:

  • human comparison 비용
  • preference subjectivity
  • reward hacking
  • length bias
  • style bias

입니다.

반대로 RLVR는 reliable하지만 reward가 sparse합니다.

또 다른 접근인 process supervision은 reasoning 과정의 step별 reward를 제공합니다.

예: r1,r2,…,rTr_1,r_2,\ldots,r_T

하지만 step-level label 자체를 만들기가 비쌉니다.

따라서 저자들이 보는 spectrum은 대략 다음과 같습니다.

방법Reward signal장점단점
RLVRexact verifier정확, 저비용적용 domain 제한
RLHF/RMpreference범용적bias/reward hacking
Process RMstep-wise correctnessdense supervisionannotation 비용
RaRstructured rubricdense + interpretablerubric 품질 의존

RaR는 이 중간 지점을 노립니다.


4. 방법론: Formalization

입력 prompt: x

현재 policy: πθ\pi_\theta

에서 response를 sample: y^∼πθ(⋅|x).\hat y\sim\pi_\theta(\cdot|x).

각 prompt x마다 k개의 rubric criterion이 존재합니다.

ℛx={(wj,cj)}j=1k.\mathcal R_x = \{(w_j,c_j)\}_{j=1}^{k}.

여기서 wj∈ℝw_j\in\mathbb R 는 criterion importance이고, cj(x,y^)∈{0,1}c_j(x,\hat y)\in\{0,1\}는 criterion j를 response가 만족했는지 나타냅니다.


5. Reward Aggregation 방법

논문에서 가장 중요한 methodological comparison입니다.

두 가지가 있습니다.

  1. Explicit aggregation
  2. Implicit aggregation

5.1 Explicit Aggregation

각 criterion을 독립적으로 judge합니다.

그 후 weighted average: r(x,y^)=∑j=1kwjcj(x,y^)∑j=1kwj\boxed{ r(x,\hat y) = \frac{ \sum_{j=1}^{k}w_j c_j(x,\hat y) }{ \sum_{j=1}^{k}w_j } } 을 사용합니다.

예를 들어 rubric이 다음과 같다고 합시다.

Criterionweightsatisfy
Correct diagnosis1.01
Explain key evidence0.71
Optional detail0.30
Avoid pitfall0.91

그러면 r=1.0(1)+0.7(1)+0.3(0)+0.9(1)1.0+0.7+0.3+0.9=2.62.9≈0.897.r= \frac{ 1.0(1)+0.7(1)+0.3(0)+0.9(1) }{ 1.0+0.7+0.3+0.9 } = \frac{2.6}{2.9} \approx0.897.

장점

Criterion별로 무엇을 만족했는지를 알 수 있으므로 interpretability가 높습니다.

단점

wjw_j를 사람이 정해야 합니다.

실험에서는: Essential=1.0Important=0.7Optional=0.3Pitfall=0.9\begin{aligned} \text{Essential}&=1.0\\ \text{Important}&=0.7\\ \text{Optional}&=0.3\\ \text{Pitfall}&=0.9 \end{aligned} 을 사용합니다.

특히 Pitfall criterion은 실제로

“misinformation을 피한다”

처럼 positive wording으로 바꾸므로 만족하면 positive reward가 들어갑니다.


5.2 Implicit Aggregation

여기가 실험상 가장 좋은 방법입니다.

criterion별 cjc_j를 명시적으로 계산해서 합산하지 않습니다.

대신 LLM judge에 (x,y^,ℛx)(x,\hat y,\mathcal R_x) 전체를 넣고 rimplicit=fϕ(x,y^,{dj}j=1k)\boxed{ r_{\text{implicit}} = f_\phi(x,\hat y,\{d_j\}_{j=1}^k) }를 직접 계산합니다.

judge에게는 대략:

prompt와 response와 rubrics가 주어졌을 때 모든 rubric을 종합적으로 고려하여 response quality를 1–10으로 평가하라.

라고 요청합니다.

그 뒤 1,…,10→[0,1]1,\ldots,10 \rightarrow [0,1] 로 normalize합니다.

중요한 차이

Explicit: judge each criterion→manually aggregate\text{judge each criterion} \rightarrow \text{manually aggregate}

Implicit: judge all criteria jointly→one holistic score\text{judge all criteria jointly} \rightarrow \text{one holistic score}

입니다.

RaR-Implicit는 사람이 wjw_j를 수동 tuning할 필요가 없습니다.


6. 왜 이것이 RLVR의 generalization인가?

논문의 흥미로운 이론적 framing입니다.

일반 RaR: r(x,y^)=∑jwjcj(x,y^)∑jwjr(x,\hat y)= \frac{\sum_jw_jc_j(x,\hat y)} {\sum_j w_j}

에서 k=1,w1=1k=1, w_1=1 이고 c1(x,y^)=match⁡(y,y^)c_1(x,\hat y) = \operatorname{match}(y,\hat y) 라고 놓으면 r(x,y^)=match⁡(y,y^).r(x,\hat y) = \operatorname{match}(y,\hat y).

즉 conventional RLVR가 그대로 됩니다.

따라서 저자들은: RLVR⊂Rubric-guided RL\boxed{ \text{RLVR}\subset\text{Rubric-guided RL} }이라고 해석합니다.

RaR는 single binary correctness를 c1c_1 에서 (c1,c2,…,ck)(c_1,c_2,\ldots,c_k)로 일반화한 것입니다.

이 framing은 이 논문의 conceptual contribution 중 상당히 좋은 부분이라고 봅니다.


7. Rubric Generation

실제로는 rubric 품질이 방법 전체의 성패를 결정합니다.

논문은 좋은 rubric이 가져야 할 네 가지 desiderata를 정의합니다.


7.1 Expert Grounding

criterion은 domain expert가 중요하게 보는:

  • essential fact
  • reasoning step
  • conclusion

을 포함해야 합니다.

하지만 의료 전문가가 모든 20K 문제에 rubric을 쓰는 것은 비쌉니다.

그래서 Reference answer→strong LLM→Rubric\boxed{\text{Reference answer}\rightarrow\text{strong LLM}\rightarrow\text{Rubric}}을 사용합니다.

즉 reference answer를 expert guidance의 proxy로 봅니다.


7.2 Comprehensive Coverage

단순 factual correctness뿐 아니라

  • accuracy
  • logical coherence
  • completeness
  • style
  • safety

등을 포함해야 합니다.

그리고 Pitfall criterion도 둡니다.

예:

“common misdiagnosis를 피해야 한다.”


7.3 Criterion Importance

모든 criterion이 동일하게 중요하지 않습니다.

예: factual correctness>stylistic clarity.\text{factual correctness} > \text{stylistic clarity}.

따라서 importance information을 붙입니다.

논문에서는:

  • Essential
  • Important
  • Optional
  • Pitfall

을 사용합니다.


7.4 Self-contained criterion

각 criterion은 judge가 다른 정보를 다시 찾아보지 않아도 평가할 수 있어야 합니다.

예를 들어 나쁜 criterion은:

“적절한 치료법을 제시한다.”

입니다.

무엇이 적절한지 criterion 자체로는 모릅니다.

좋은 criterion은:

“The response recommends approximately 150 mEq sodium bicarbonate during the first four hours.”

처럼 작성합니다.

이렇게 해야 작은 judge도 criterion 자체만 보고 cj(x,y)c_j(x,y)를 평가할 수 있습니다.

네 가지 설계 원칙과 7–20개의 prompt-specific criteria 생성 과정이 논문에 명시되어 있습니다.


8. 실제 의료 Rubric 예

논문의 sodium bicarbonate 예가 이 방법을 아주 잘 보여줍니다.

문제:

65 kg 환자, pH=7.05, base deficit=-40. 처음 4시간 동안 bicarbonate를 얼마나 투여할 것인가?

Reference answer에서는 40×65×0.3=780 mEq40\times65\times0.3=780\text{ mEq} 를 계산한 뒤, safety를 위해 처음에는 약 150 mEq를 투여한다고 합니다.

여기에 생성되는 rubric은 다음과 비슷합니다.

Essential

  1. 공식

Base Deficit×Weight×0.3\text{Base Deficit}\times \text{Weight}\times0.3

을 올바르게 사용.

  1. 약 150 mEq의 초기 dose 권고.

Important

  1. partial correction의 이유 설명.
  2. 780 mEq 계산 과정 설명.
  3. patient data를 올바르게 사용.

Optional

  1. base deficit이 severe metabolic acidosis를 의미한다는 설명.

Pitfall

  1. rapid overcorrection 위험을 무시하지 않음.

즉 하나의 reference answer를 단순히 generated answer≈reference answer?\text{generated answer}\approx\text{reference answer}?로 비교하지 않고,reference answer→atomic requirements\boxed{\text{reference answer}\rightarrow\text{atomic requirements}}로 분해합니다.

이 점이 Reference-Likert보다 효과적인 이유입니다.


9. Rubric dataset

9.1 RaR-Medicine

약 20K medical prompts입니다.

source:

  • medical-o1-reasoning-SFT
  • NaturalReasoning
  • SCP-116K
  • GeneralThought-430K

Rubric은 GPT-4o로 생성합니다.

실제 전체 예제 수는: 20,166

평균 rubric 수: 7.5

평균 question length: 45.0 words.45.0\text{ words}.


9.2 RaR-Science

약 20,625

examples.

GPQA-Diamond의 분야와 맞도록:

  • chemistry
  • physics
  • biology

등의 science reasoning question을 구성합니다.

Rubric은 o3-mini로 생성합니다.

평균 rubric 수도 약 7.5 입니다.


10. Policy learning: GRPO

Base model은 Qwen2.5-7B 입니다.

각 prompt q 마다: K=16개 response를 현재 policy에서 sampling합니다.

y1,…,y16∼πθ(⋅|q).y_1,\ldots,y_{16} \sim \pi_\theta(\cdot|q).

Sampling temperature: T=1.0

context/max length: 3584.

judge는 기본적으로: GPT-4o-mini 입니다.

각 rollout에 대해 Ri=RubricJudge(q,yi,ℛq)R_i = \text{RubricJudge}(q,y_i,\mathcal R_q)를 계산합니다.

그 다음 GRPO를 통해 같은 prompt 내 response들의 상대적 reward로 policy를 update합니다.

논문의 training 설정은: rollouts/prompt=16effective batch size=96LR=5×10−6warmup ratio=0.1training steps=300\begin{aligned} \text{rollouts/prompt}&=16\\ \text{effective batch size}&=96\\ \text{LR}&=5\times10^{-6}\\ \text{warmup ratio}&=0.1\\ \text{training steps}&=300 \end{aligned}

이고 8×H100 한 노드에서 학습합니다.


GRPO 관점에서 보면

논문 자체는 GRPO objective를 새롭게 제안한 것이 아니라 GRPO에 들어가는 reward RiR_i를 바꾼 연구입니다.

일반적인 GRPO 관점에서는 동일 prompt의 reward들에 대해 대략 Ai=Ri−mean⁡(R1,…,RK)std⁡(R1,…,RK)+ϵA_i = \frac{ R_i-\operatorname{mean}(R_1,\ldots,R_K) }{ \operatorname{std}(R_1,\ldots,R_K)+\epsilon } 형태의 group-relative advantage를 만들고 policy를 개선합니다.

따라서 RaR의 핵심은 optimizer가 아니라 Ri=무엇으로 정의하는가?\boxed{ R_i=\text{무엇으로 정의하는가?} } 입니다.

기존 RLVR: Ri=correct / incorrectR_i=\text{correct / incorrect}

Direct-Likert: Ri=judge의 막연한 1–10 평가R_i=\text{judge의 막연한 1–10 평가}

RaR: Ri=prompt-specific rubric을 사용한 structured evaluation\boxed{ R_i=\text{prompt-specific rubric을 사용한 structured evaluation} } 입니다.


11. Baseline

비교 구성이 매우 중요합니다.

(1) Qwen2.5-7B

base model.

(2) Qwen2.5-7B-Instruct

instruction-tuned model.


(3) Direct-Likert

Judge에게

“이 response가 얼마나 좋은지 1–10점으로 평가하라.”

만 묻습니다. Rdirect=Norm⁡(Likert⁡(q,y)).R_{\mathrm{direct}} = \operatorname{Norm} ( \operatorname{Likert}(q,y) ).

즉 rubric도 reference도 없습니다.


11.1 Reference-Likert

Judge에게 reference answer까지 제공합니다.

Rref=Norm⁡(Likert⁡(q,y,y∗)).R_{\mathrm{ref}} = \operatorname{Norm} ( \operatorname{Likert} (q,y,y^*) ).

RaR와 비교하기 위한 상당히 강한 baseline입니다.


11.2 RaR-Predefined

모든 문제에 동일한 generic rubric을 사용합니다.

예:

  • factually correct
  • complete
  • concise
  • helpful

따라서 instance-specificity가 없습니다.


11.3 RaR-Explicit

Prompt-specific rubric + explicit weighted sum.


11.4 RaR-Implicit

Prompt-specific rubric + holistic LLM aggregation.

최종적으로 이 방법이 가장 좋습니다.


12. 평가 1: HealthBench

HealthBench는 5,000개의 realistic clinical conversation을 포함하고 physician-authored rubric으로 평가합니다.

평가 axis:

  • Communication quality
  • Instruction following
  • Accuracy
  • Context awareness
  • Completeness

Main overall result:

MethodHealthBench overall
Qwen2.5-7B7.7
Qwen2.5-7B-Instruct22.7
Direct-Likert25.5
Reference-Likert28.9
RaR-Predefined12.5
RaR-Explicit29.7
RaR-Implicit31.2

중요한 비교

Direct-Likert: 25.5

RaR-Implicit: 31.2

relative improvement: 31.2−25.525.5≈22.4%.\frac{31.2-25.5}{25.5} \approx22.4\%.

본문이 언급하는 “up to 31%”는 세부 evaluation setting/axis의 최대 relative improvement를 포함한 표현입니다.

더 중요한 것은 Reference-Likert도 넘어선다는 점입니다.

31.2>28.9.

즉 같은 reference 정보가 있다 하더라도Reference answer를 그대로 judge에게 주는 것<Reference를 rubric으로 구조화해 주는 것\boxed{ \text{Reference answer를 그대로 judge에게 주는 것} < \text{Reference를 rubric으로 구조화해 주는 것} }

이라는 결과입니다.


13. RaR-Predefined가 왜 실패하는가?

이 결과가 상당히 중요합니다.

RaR-Predefined: 12.5 로 매우 낮습니다.

심지어 Direct-Likert: 25.5 보다 훨씬 나쁩니다.

이는 단순히

“rubric을 쓰면 좋다”

가 아니라 𝐢𝐧𝐬𝐭𝐚𝐧𝐜𝐞-𝐬𝐩𝐞𝐜𝐢𝐟𝐢𝐜 𝐫𝐮𝐛𝐫𝐢𝐜이어야 한다\boxed{ \textbf{instance-specific rubric이어야 한다} }

는 것을 의미합니다.

Generic criteria:

  • correct
  • concise
  • complete

만으로는 어떤 사실이 정확해야 하는지, 어떤 failure를 피해야 하는지 reward model에 알려주지 못합니다.

저자들도 generic rubric이 prompt-specific requirement와 failure mode를 놓쳐 reward misalignment를 유발한다고 분석합니다.


14. 평가 2: GPQA-Diamond

이 부분이 논문의 generalization claim에 중요합니다.

GPQA-Diamond는 multiple-choice science reasoning benchmark입니다.

즉 evaluation 시에는 rubric이 필요 없습니다.

결과:

MethodGPQA-Diamond
Qwen2.5-7B31.7
Qwen2.5-7B-Instruct35.0
Direct-Likert34.8
Reference-Likert36.5
RaR-Predefined31.7
RaR-Explicit36.9
RaR-Implicit37.6

Direct-Likert 대비: 34.8→37.6.34.8\rightarrow37.6.

대략 37.6−34.834.8≈8.0%\frac{37.6-34.8}{34.8} \approx8.0\% 수준이며 논문에서는 약 7% relative gain으로 요약합니다.


15. 왜 GPQA 결과가 중요한가?

가능한 비판은:

“rubric reward로 학습했으니 rubric으로 평가하는 HealthBench에서 잘 나오는 것은 당연하지 않은가?”

입니다.

GPQA-Diamond는 이 문제를 완화합니다.

Training: Rubric reward

Evaluation: multiple-choice accuracy.

그럼에도 성능이 증가했습니다.

따라서 저자들의 주장은: 모델이 단순히 rubric evaluator를 gaming한 것이 아니라, 실제 reasoning capability도 어느 정도 향상되었다

입니다.

물론 이것만으로 reward hacking 가능성이 완전히 제거됐다고 보기는 어렵습니다.


16. RaR-Implicit vs RaR-Explicit

결과는 일관되게: Implicit>Explicit.\text{Implicit}>\text{Explicit}.

HealthBench: 31.2>29.7

GPQA: 37.6>36.9.

저자 해석은:

Explicit

장점:

  • criterion별 score 확인 가능
  • weight 조정 가능
  • interpretability 우수

단점:

  • hand-tuned weight가 brittle
  • criterion 간 상호작용을 표현하기 어려움

Implicit

장점:

  • judge가 criterion interaction을 종합적으로 판단
  • hand-designed weighting 불필요

즉 예를 들어

“진단은 틀렸지만 설명 스타일은 좋다”

라는 response를 단순 weighted sum보다 LLM이 holistic하게 낮게 판단할 수 있다는 것입니다.

논문도 Explicit의 고정 weight는 통제성과 해석성은 높지만 brittle하며, Implicit가 전반적으로 가장 좋다고 결론냅니다.


17. Human preference alignment 실험

또 하나 흥미로운 결과입니다.

약 3,000개의 HealthBench prompt를 사용하여:

  • preferred: practitioner-approved answer
  • rejected: controlled perturbation으로 열화시킨 response

pair를 만듭니다.

Judge에게 두 response를 평가하게 해서 Pairwise Accuracy=P(s(y+)>s(y−))\text{Pairwise Accuracy} = P( s(y^+)>s(y^-) ) 를 측정합니다.

비교:

Direct Likert

(q,y)→score(q,y) \rightarrow score

Rubric-guided

(q,y,ℛq)→score.(q,y,\mathcal R_q) \rightarrow score.

결과적으로 모든 judge scale에서 rubric 제공이 human preference와의 alignment를 높였습니다.

특히 작은 judge에서 개선폭이 큽니다.


18. 작은 Judge가 특히 이득을 보는 이유

저자들의 핵심 해석입니다.

작은 LLM에게 그냥

“이 의료 답변에 1~10점을 줘라.”

라고 하면 작은 모델이 내부적으로 다음을 모두 추론해야 합니다. 무엇이 중요한가?어떤 오류가 심각한가?무엇이 빠졌는가?안전성 문제는 무엇인가?\begin{aligned} &\text{무엇이 중요한가?}\\ &\text{어떤 오류가 심각한가?}\\ &\text{무엇이 빠졌는가?}\\ &\text{안전성 문제는 무엇인가?} \end{aligned}

반면 rubric을 주면:

  1. diagnosis X를 언급했는가?
  2. sign Y와 diagnosis를 연결했는가?
  3. treatment Z를 권고했는가?
  4. dangerous recommendation W를 피했는가?

처럼 evaluation problem을 분해합니다.

즉 rubric이 일종의 evaluation scaffold\boxed{\text{evaluation scaffold}}역할을 합니다.

Appendix의 judge-scale training 결과도 이를 보여줍니다.

JudgeRaR-ImplicitDirect-Likert
GPT-4o-mini27.925.3
Qwen-32B26.225.4
Qwen-14B25.024.9
Qwen-7B26.722.0

Qwen-7B에서는: 22.0→26.722.0\rightarrow26.7

즉 +4.7 points 향상입니다.

논문에서는 rubric reward 사용 시 judge scale에 따른 policy 성능 범위도 더 좁아진다고 해석합니다.


19. Ablation 1: Human vs Synthetic Rubric

HealthBench-1k에서 가장 흥미로운 ablation입니다.

TrainingOverall
Expert-Answer-SFT20.4
Simple-Likert23.9
Reference-Likert31.7
RaR-Implicit-Synthetic-NoRef32.0
RaR-Implicit-Synthetic35.9
RaR-Implicit-Human34.8

여기서:

  • Synthetic-NoRef: LLM이 reference 없이 rubric 생성
  • Synthetic: reference answer를 보고 rubric 생성
  • Human: 인간이 작성한 rubric

입니다.

가장 놀라운 결과: 35.9synthetic+reference>34.8human.35.9_{\text{synthetic+reference}} > 34.8_{\text{human}}.

차이는 크지 않으므로 “synthetic이 human보다 본질적으로 우월하다”고 해석하면 안 됩니다.

보다 적절한 결론은: strong LLM + high-quality reference로 만든 rubric이 human-authored rubric에 근접하는 수준의 supervision을 제공할 수 있다

입니다.


20. 왜 reference answer가 중요한가?

Synthetic rubric, no reference: 32.0

Synthetic rubric + reference: 35.9.

즉 LLM 자체가 아무리 강해도 문제만 보고

“좋은 답이 갖추어야 할 criterion을 생성하라.”

라고 하는 것보다, 정답/reference를 제공한 뒤 criterion을 생성하는 것이 훨씬 좋습니다.

이것은 RaR의 scalability에 중요한 함의를 줍니다.

사람이 직접 rubric 10개를 쓰는 대신: Human/expert answer –> LLM rubric expansion

을 수행할 수 있습니다.

저자도 reference-guided synthetic rubric이 reference-free rubric보다 일관되게 강하며, high-stakes domain에서는 expert signal이 중요하다고 강조합니다.


21. Ablation 2: Rubric의 어떤 구성 요소가 중요한가?

결과:

SettingHealthBench-1k
Essential-only34.9
No categorical labels38.8
No pitfall37.2
All rubrics37.2

이 결과는 꽤 흥미롭습니다.


21.1 Essential-only는 나쁨

34.9

즉 정답의 핵심만 검사하면 충분하지 않습니다.

다양한 criteria:

  • correctness
  • explanation
  • completeness
  • safety
  • etc.

를 모두 사용하는 것이 좋은 learning signal을 만듭니다.


21.2 Importance label은 의외로 별로 중요하지 않음

놀랍게도: No categorical labels=38.8가 All=37.2 보다 오히려 높습니다.

따라서

  • Essential
  • Important
  • Optional

이라는 hand-designed importance 정보가 필수는 아닙니다.

이 결과는 특히 RaR-Implicit와 잘 연결됩니다.

LLM judge가 criterion 내용을 읽고 중요도를 스스로 어느 정도 infer할 수 있기 때문입니다.


21.3 Pitfall도 효과가 명확하지 않음

No Pitfall=37.2 , All=37.2.

즉 synthetic pitfall criteria가 큰 도움을 주지 못했습니다.

저자들의 설명은:

어떤 오류가 실제 모델에게 흔하고 중요한 failure mode인지 예측하는 것이 어렵기 때문이다.

입니다.

이는 중요한 future-work 포인트입니다.


22. Ablation 3: Rubric generator가 강할수록 좋은가?

reference를 주지 않은 조건에서 비교합니다.

Rubric generatorHealthBench
GPT-4o34.2
GPT-4o-mini32.7
o3-mini32.4
Qwen-72B32.7
Qwen-32B31.1
Qwen-7B31.9
o3-mini + reference35.9

두 가지 결론이 나옵니다.

결론 1

일반적으로 stronger rubric generator가 유리합니다.

하지만 parameter scale과 단순 monotonic relationship은 아닙니다.

예: GPT-4o-mini≈Qwen-72B\text{GPT-4o-mini}\approx\text{Qwen-72B} 입니다.

따라서 instruction alignment + reasoning ability가 중요합니다.

결론 2

더 중요한 것은 reference입니다.

o3-mini + reference=35.9\text{o3-mini + reference}=35.9 가 GPT-4o no reference=34.2 보다 높습니다.

즉, expert grounding > rubric generator의 단순 모델 크기

라는 결과로 볼 수 있습니다.


23. 전체 실험 결과를 한 장으로 요약하면

A. Instance-specific rubric이 중요

RaR-Implicit=31.2 versus RaR-Predefined=12.5.


B. Structured reference가 raw reference보다 좋음

RaR-Implicit=31.2 > Reference-Likert=28.9.

즉, y∗→LLM judgey^* \rightarrow \text{LLM judge}

보다 y∗→{c1,…,ck}→LLM judgey^* \rightarrow \{c_1,\ldots,c_k\} \rightarrow \text{LLM judge} 가 좋습니다.


C. Implicit aggregation이 가장 강함

RaR-Implicit > RaR-Explicit.


D. Rubric learning이 non-rubric evaluation으로 transfer

GPQA:\34.8→37.6.34.8 \rightarrow 37.6.


E. Smaller judge가 가장 큰 이득

특히 Qwen-7B judge: 22.0→26.7.22.0 \rightarrow26.7.

Rubric이 evaluator의 reasoning burden을 줄여줍니다.


F. Rubric generator보다 expert grounding이 중요

o3-mini + reference>GPT-4o without reference.\text{o3-mini + reference} > \text{GPT-4o without reference}.


24. 이 논문의 방법을 보는 가장 중요한 관점

이 논문을 단순히

“LLM judge에게 rubric을 줬더니 성능이 올라갔다.”

라고 보면 연구적 의미를 과소평가하게 됩니다.

좀 더 본질적으로 보면 reward function specification의 문제입니다.

기존 RLHF: Human preference→Reward model\boxed{ \text{Human preference} \rightarrow \text{Reward model} }

RLVR:Ground-truth verifier→Reward\boxed{ \text{Ground-truth verifier} \rightarrow \text{Reward} }

RaR: Expert intent→atomic criteria→LLM-based structured verifier→Reward\boxed{ \text{Expert intent} \rightarrow \text{atomic criteria} \rightarrow \text{LLM-based structured verifier} \rightarrow \text{Reward} } 입니다.

즉 어려운 것은 R(x,y) 라는 복잡한 reward function을 학습하는 것이 아니라,R(x,y)=F(c1(x,y),…,ck(x,y))R(x,y) = F(c_1(x,y),\ldots,c_k(x,y)) 로 분해하여 specification하는 것이라는 관점입니다.


25. 저자들이 제시하는 Future Work

논문의 limitation/future work는 세 가지입니다.

① 더 넓은 domain

현재:

  • medicine
  • science

뿐입니다.

향후:

  • dialogue
  • tool use
  • agentic tasks

등을 봐야 합니다.

② Dynamic / learned rubric weighting

현재 explicit: wj=fixed.w_j=\text{fixed}.

향후: wj=wj(x,t)w_j=w_j(x,t) 처럼 학습하거나 RL 단계에 따라 바꿀 수 있습니다.

예:

초기: wcorrectness≫wstylew_{\text{correctness}}\gg w_{\text{style}}

후기: wsafety,wstyle↑.w_{\text{safety}},w_{\text{style}}\uparrow.

즉 curriculum reward입니다.

③ Specialized evaluator

현재는 off-the-shelf LLM judge를 사용합니다.

향후:

  • reward reasoning model
  • specialized rubric verifier
  • generative reward model

을 사용할 수 있습니다.



게시됨

카테고리

,

작성자

댓글

답글 남기기

이메일 주소는 공개되지 않습니다. 필수 필드는 *로 표시됩니다