Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
對比解碼緩解 LLM 評審中的分數範圍偏差
原始論文:Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
作者:Yoshinari Fujinuma
機構:Cantina Labs
arXiv ID:2510.18196v1
日期:2025-10-21
標籤:LLM-as-a-Judge Evaluation Decoding Bias Summarization NLP
目錄
- 摘要
- 1 引言
- 2 相關工作
- 3 用於緩解 LLM 評審偏差的對比解碼
- 4 實驗
- 5 結論
- 局限性
- 附錄 A 評審提示詞
- 附錄 B 超參數
- 附錄 C 相關性與一致性結果
- 附錄 D 模型大小與預算
- 附錄 E AI 輔助工具使用說明
- 附錄 F 潛在風險
- 參考文獻
- 術語對照表
摘要
大型語言模型 (Large Language Models, LLM) 常被用作各種應用中的評估器,但其結果的可靠性仍是一大挑戰。其中一個挑戰是將 LLM 用作直接評估 (Direct Assessment) 的評審,即在沒有任何參考的情況下從指定範圍內分配分數。我們首先表明,這一挑戰源於 LLM 評審輸出與分數範圍偏差 (Score Range Bias) 的關聯,即 LLM 評審輸出對預定義的分數範圍高度敏感,阻礙了對最佳分數範圍的搜索。我們還發現,來自同一模型家族的模型之間存在類似的偏差。隨後,我們透過對比解碼 (Contrastive Decoding) 來緩解這種偏差,在不同分數範圍上的 Spearman 相關性與人類判斷相比,平均實現了高達 11.3% 的相對提升。
1 引言
大型語言模型評審已成為評估生態系統中不可或缺的組成部分 (Lin et al., 2022; Chiang and Lee, 2023; Bubeck et al., 2023)。從直接評估——評審透過分配分數來評估個別輸出——(Liu et al., 2023) 到成對比較 (Pairwise Comparison)——評審比較兩個輸出並確定哪個更優——(Zheng et al., 2023; Ye et al., 2024),使用 LLM 作為評審越來越多地被部署,以提供跨多樣化任務的自動、可擴展且具成本效益的評估。然而,此類評估的可靠性面臨重大挑戰,特別是當模型評估自身輸出 (Zheng et al., 2023) 或來自同一模型家族的輸出 (Goel et al., 2025) 時。
這些偏差限制了可以可靠地用作 LLM 評審的模型集合。
但在使用 LLM 作為評審時,是否還隱藏著其他偏差?

圖 1:2-4 分數範圍中分數範圍偏差的概述,以及對比解碼 (Contrastive Decoding) 如何透過消除同一模型家族中的類似偏差來緩解它。
在本研究中,我們揭示了 LLM 評審輸出中的另一種偏差,即分數範圍偏差,其中 LLM 評審輸出對分數範圍的移動非常敏感,這一現象的動機來自先前的發現:LLM 在簡單算術任務中表現不佳 (Nogueira et al., 2021; Gambardella et al., 2024)。
在識別出這些偏差後,我們還透過將近期的對比解碼研究 (Li et al., 2023; O'Brien and Lewis, 2023) 與家族增強偏差 (Family-Enhancement Bias) (Goel et al., 2025) 聯繫起來,探索了一種緩解策略,旨在消除同一模型家族中編碼的類似分數範圍偏差。
我們的主要貢獻總結如下:
- 我們首先表明 LLM 評審存在分數範圍偏差——這是在不同模型大小和家族 (Llama-3 和 Qwen-2.5) 中觀察到的偏差,發生在直接評估的評審任務中。
- 我們接著證明,對比解碼基於同一模型家族中存在類似分數範圍偏差的觀察,成功地緩解了這些偏差。
2 相關工作
我們現在回顧 LLM 評審的相關工作,聚焦於評審任務及其偏差。
LLM 評審任務
LLM 評審任務分為兩類:直接或逐點評估 (Pointwise Assessment) (Jones et al., 2024; Li et al., 2024; Zhu et al., 2025) 和成對評估 (Zheng et al., 2023; Ye et al., 2024)。直接評估 (Liu et al., 2023) 涉及 LLM 評審在沒有其他輸出參考的情況下為輸出分配數值評分。
在成對評估中,LLM 評審與人類偏好的相關性高於直接評估 (Liu et al., 2024),這支持了直接評估中仍存在挑戰的觀點,因此我們聚焦於直接評估的實驗。
LLM 評審偏差
LLM 評審中一個已知的偏差是自我增強偏差 (Self-Enhancement Bias)——傾向於偏好自身輸出 (Liu et al., 2023; Zheng et al., 2023; Ye et al., 2024),即使在 GPT-4 等專有模型中也存在 (Wataoka et al., 2024)。
在自我增強偏差之外,Goel et al. (2025) 報告了家族增強偏差,即模型偏好來自同一模型家族的輸出。
假設這些家族偏差是雙向的,我們旨在使用對比解碼來消除它們。
3 用於緩解 LLM 評審偏差的對比解碼
對比解碼 (Li et al., 2023) 透過使用兩個模型來修改模型輸出:一個主模型 (Main Model) 和一個輔助模型 (Assistant Model)。
給定主模型的下一個 token 機率 $p_{\text{main}}$ 和輔助模型的 $p_{\text{asst}}$,最終調整後的分數透過從 $p_{\text{main}}$ 中減去加權的 $p_{\text{asst}}$ 來計算,即:
$$\log p_{\text{main}} - \lambda \log p_{\text{asst}} \tag{1}$$
其中 $\lambda \in \mathbb{R}$ 是控制輔助模型幅度的超參數 (Hyperparameter),token $i$ 的 logit $e_i$ 由溫度 $t > 0$ 控制,即:
$$p_{\text{Asst}} = \frac{e_i / t}{\sum_j e_j / t}$$
與 Li et al. (2023) 相比,我們的一個不同之處是加入了 $\lambda$,以進一步對齊兩個模型之間的 logit 分佈,這受到了我們在 §4.2 中分析的啟發。
4 實驗
我們聚焦於摘要 (Summarization) 的直接評估,因為先前的研究報告 LLM 評審在此方面表現不足 (Ye et al., 2024),且摘要是常用的評估場景 (Panickssery et al., 2024)。
4.1 實驗設定
任務與指標
我們聚焦於摘要任務,LLM 評審通常在此場景中使用 (Liu et al., 2023; Panickssery et al., 2024)。人類標註之間的相關性使用三個指標衡量:Pearson 相關性、Spearman 相關性和 Kendall 相關性。
分數量表與範圍
我們使用 5 點 Likert 量表 (Likert, 1932) 在不同的分數範圍上 (0-4, 1-5, 2-6, 3-7)。【譯註:作者在 7 處停止,靈感來自 Likert (1932) 所示的 5 點 (1-5) 和 7 點 (1-7) 結果之間的高相關性。】
如果輸出分數解析失敗,我們遵循 Liu et al. (2023) 的做法將其設為最低分;如果解析的分數超過最大值,我們將其截斷 (Clamp) 到範圍內的最高分。
模型
我們在兩個模型家族上進行實驗。【譯註:作者將專門為評審任務微調的模型(如 Prometheus)留作未來工作,因為這些模型沒有多個模型大小可用於對比解碼,且這些模型是針對 1-5 分數範圍進行微調的。】
對於 Llama-3 家族 (Grattafiori et al., 2024),我們使用 Llama-3.1-8B-Instruct 作為主模型,Llama-3.2-3B-Instruct 和 Llama-3.2-1B-Instruct 作為輔助模型;對於 Qwen2.5 家族 (Qwen et al., 2025),我們使用 Qwen-2.5-14B-Instruct 和 Qwen-2.5-7B-Instruct 作為主模型,Qwen-2.5-3B-Instruct 作為輔助模型。評審提示詞見附錄 A。

圖 2(a):Llama-3 家族結果

圖 2(b):Qwen2.5 家族結果
圖 2:2-6 分數範圍中使用貪婪解碼 (Greedy Decoding)、對比解碼和人類標註的連貫性 (Coherence) 分數分佈。Llama 8B 和 3B 模型的貪婪解碼輸出高度偏向輸出分數 4,而 Qwen2.5 3B 和 7B 模型則輸出分數 2,顯示這些模型中編碼了類似的偏差。
資料集
我們使用 SummEval (Fabbri et al., 2021),這是一個摘要基準測試,也被 Liu et al. (2023) 使用,包含 100 篇新聞文章,每篇文章關聯 16 個摘要及人類標註分數,共計 1600 個摘要。我們使用 10% 的新聞文章作為留出的開發集。
我們現在首先透過評估在相同 5 點量表但不同分數範圍上與人類標註的相關性來揭示 LLM 評審的分數範圍偏差,然後進行緩解實驗。

圖 3(a):Qwen2.5 3B logit 分佈

圖 3(b):Qwen2.5 7B logit 分佈

圖 3(c):Qwen2.5 14B logit 分佈
圖 3:0-4 分數範圍中第一個輸出 token 的 logit 分佈。對於 3B 和 7B 模型,分數 2 的 logit 最高,而隨著模型大小從 3B 增加到 14B,分數 3 的 logit 逐漸升高並更接近人類判斷。
表 1:Llama-3 家族在摘要連貫性上與人類標註的相關性結果
| 模型 | 範圍 | Pear. | Spear. | Kend. |
|---|---|---|---|---|
| Llama 3.2-1B-Inst | 0 to 4 | 0.112 | 0.106 | 0.092 |
| Llama 3.2-3B-Inst | 0 to 4 | 0.141 | 0.146 | 0.126 |
| Llama 3.1-8B-Inst | 0 to 4 | 0.381 | 0.367 | 0.321 |
| Contrastive (8B-1B) | 0 to 4 | 0.386 | 0.369 | 0.324 |
| Contrastive (8B-3B) | 0 to 4 | 0.380 | 0.361 | 0.316 |
| Llama 3.2-1B-Inst | 1 to 5 | 0.071 | 0.069 | 0.060 |
| Llama 3.2-3B-Inst | 1 to 5 | 0.205 | 0.217 | 0.189 |
| Llama 3.1-8B-Inst | 1 to 5 | 0.347 | 0.338 | 0.293 |
| Contrastive (8B-1B) | 1 to 5 | 0.365 | 0.358 | 0.311 |
| Contrastive (8B-3B) | 1 to 5 | 0.340 | 0.336 | 0.293 |
| Llama 3.2-1B-Inst | 2 to 6 | 0.000 | 0.000 | 0.000 |
| Llama 3.2-3B-Inst | 2 to 6 | 0.168 | 0.167 | 0.148 |
| Llama 3.1-8B-Inst | 2 to 6 | 0.270 | 0.257 | 0.226 |
| Contrastive (8B-1B) | 2 to 6 | 0.310 | 0.302 | 0.264 |
| Contrastive (8B-3B) | 2 to 6 | 0.302 | 0.298 | 0.258 |
| Llama 3.2-1B-Inst | 3 to 7 | -0.104 | -0.128 | -0.105 |
| Llama 3.2-3B-Inst | 3 to 7 | 0.033 | 0.045 | 0.040 |
| Llama 3.1-8B-Inst | 3 to 7 | 0.386 | 0.372 | 0.320 |
| Contrastive (8B-1B) | 3 to 7 | 0.383 | 0.378 | 0.326 |
| Contrastive (8B-3B) | 3 to 7 | 0.393 | 0.375 | 0.324 |
| 所有分數範圍的平均值 | ||||
| Llama 3.2-1B-Inst | 0.020 | 0.012 | 0.012 | |
| Llama 3.2-3B-Inst | 0.137 | 0.144 | 0.126 | |
| Llama 3.1-8B-Inst | 0.346 | 0.334 | 0.290 | |
| Contrastive (8B-1B) | 0.361 | 0.352 | 0.306 | |
| Contrastive (8B-3B) | 0.354 | 0.343 | 0.298 |
各分數範圍內最大相關性以底線標示,所有範圍中的最大值以斜體標示,平均值中的最大值以粗體標示。最大改善出現在 2-6 分數範圍。
表 2:Qwen-2.5 家族在摘要連貫性上與人類標註的相關性結果
| 模型 | 範圍 | Pear. | Spear. | Kend. |
|---|---|---|---|---|
| Qwen2.5-3B-Inst | 0 to 4 | -0.241 | -0.221 | -0.193 |
| Qwen2.5-7B-Inst | 0 to 4 | 0.264 | 0.266 | 0.234 |
| Qwen2.5-14B-Inst | 0 to 4 | 0.424 | 0.428 | 0.375 |
| Contrastive (7B-3B) | 0 to 4 | 0.330 | 0.333 | 0.291 |
| Contrastive (14B-3B) | 0 to 4 | 0.440 | 0.449 | 0.394 |
| Qwen2.5-3B-Inst | 1 to 5 | -0.098 | -0.086 | -0.073 |
| Qwen2.5-7B-Inst | 1 to 5 | 0.385 | 0.385 | 0.333 |
| Qwen2.5-14B-Inst | 1 to 5 | 0.460 | 0.456 | 0.395 |
| Contrastive (7B-3B) | 1 to 5 | 0.360 | 0.351 | 0.304 |
| Contrastive (14B-3B) | 1 to 5 | 0.457 | 0.459 | 0.397 |
| Qwen2.5-3B-Inst | 2 to 6 | -0.108 | -0.086 | -0.075 |
| Qwen2.5-7B-Inst | 2 to 6 | 0.373 | 0.363 | 0.312 |
| Qwen2.5-14B-Inst | 2 to 6 | 0.292 | 0.302 | 0.262 |
| Contrastive (7B-3B) | 2 to 6 | 0.377 | 0.363 | 0.312 |
| Contrastive (14B-3B) | 2 to 6 | 0.391 | 0.412 | 0.361 |
| Qwen2.5-3B-Inst | 3 to 7 | 0.301 | 0.313 | 0.273 |
| Qwen2.5-7B-Inst | 3 to 7 | 0.368 | 0.363 | 0.315 |
| Qwen2.5-14B-Inst | 3 to 7 | 0.354 | 0.350 | 0.304 |
| Contrastive (7B-3B) | 3 to 7 | 0.355 | 0.344 | 0.297 |
| Contrastive (14B-3B) | 3 to 7 | 0.407 | 0.411 | 0.353 |
| 所有分數範圍的平均值 | ||||
| Qwen2.5-3B-Inst | -0.036 | -0.020 | -0.017 | |
| Qwen2.5-7B-Inst | 0.343 | 0.339 | 0.294 | |
| Qwen2.5-14B-Inst | 0.383 | 0.384 | 0.334 | |
| Contrastive (7B-3B) | 0.356 | 0.348 | 0.301 | |
| Contrastive (14B-3B) | 0.424 | 0.433 | 0.376 |
各分數範圍內最大相關性以底線標示,所有範圍中的最大值以斜體標示,平均值中的最大值以粗體標示。最大改善出現在 2-6 分數範圍。
4.2 揭示與緩解分數範圍偏差
同一模型家族中存在類似的分數範圍偏差
我們首先分析 2-6 分數範圍中輸出分數的分佈(圖 2)。Llama 家族模型 (3B 和 8B) 傾向於輸出分數 4(圖 2(a)),而 Qwen 2.5 家族模型傾向於輸出分數 2(圖 2(b))。透過使用對比解碼,這些對特定範圍的偏差得到緩解,使分數輸出更接近人類標註。
在分析 Qwen 家族模型在 0-4 分數範圍中第一個輸出 token 的 logit 分佈(圖 3)時,Qwen-2.5 3B、7B 和 14B 模型編碼了類似的偏差,其中分數 2 的 logit 最高,而最頻繁的人類標註是分數 3。對分數 2 的偏差隨著模型大小從 3B 擴展到 14B 而逐漸減少,但即使在 14B 模型中仍然存在。此外,每個模型中的 logit 範圍不同,例如 3B 的最大 logit $\approx 25$(圖 3(a)),7B $\approx 30$(圖 3(b)),14B $\approx 34$,這進一步激發了在公式 1 中加入 $\lambda$ 以對齊這些模型之間 logit 分佈的動機。3B 模型中對分數 2 的偏差有助於在作為輔助模型使用時減少 7B 和 14B 模型中編碼的類似偏差。
由於分數範圍偏差,使用 Llama 3B 或 7B 進行貪婪解碼導致 2-6 分數範圍中最低的相關性(表 1)。這一趨勢不僅限於 Llama-3 家族模型,在 Qwen-2.5 家族模型中也觀察到了(表 2)。聚焦於貪婪解碼,Qwen-2.5 家族和 Llama3.1-3B 顯示出更清晰的趨勢:1-5 分數範圍在實驗的分數範圍中顯示最高的相關性 (7B, 1 to 5: 0.385, 14B, 1 to 5: 0.456),而 Llama3.1-8B 是例外,3-7 分數範圍在所有分數範圍中顯示最高的相關性。
這些結果進一步引發了對將 LLM 評審應用於標準 1-5 範圍之外的擔憂。【譯註:也適用於 1-7 Likert 量表或 1-10 量表 (Ye et al., 2024)。】
對比解碼是跨不同分數範圍的穩健緩解策略
表 1 和表 2 進一步表明,對比解碼在不同分數範圍上展現出一致的相關性,解決了 LLM 評審中觀察到的分數範圍偏差。雖然使用單一模型在分數範圍移動時會遭受相關性下降,但對比解碼無論分數範圍如何,都與人類判斷保持更穩定的相關性(表 1)。這種穩健性在 2-6 範圍中尤為明顯,其中 Llama-3 家族的對比解碼達到了 0.310 的 Pearson 相關性(相比之下 Llama 3.2-3B 為 0.168,Llama 3.1-8B 為 0.270),在 Spearman 和 Kendall 相關性上也有類似的改善。在所有分數範圍的平均值上,Llama 8B 也取得了 5.1% 的相對改善 ($0.335 \rightarrow 0.352$),Qwen 14B 則取得了 11.3% 的相對改善 ($0.389 \rightarrow 0.433$)。
跨不同評分範圍的穩定性使得能夠搜索超越 1-5 範圍的最佳分數範圍(例如,在附錄 C 中 Qwen 家族的摘要相關性中,0-4 範圍顯示最佳相關性)。
輔助模型的選擇是否影響偏差緩解?
表 1 顯示,輔助模型的選擇略微影響相關性,1B 模型略優於 3B 模型。1B 輔助模型達到了 0.352 的平均 Spearman 相關性,而 3B 輔助模型為 0.343。然而,差異很小,且取決於摘要的評估維度(附錄 C)。
5 結論
在本研究中,我們分析並實驗了 LLM 作為直接評估的評審,揭示了兩個關鍵發現:首先,LLM 評審在不同模型家族和大小中展現出分數範圍偏差,傾向於偏好特定分數,而不考慮摘要的品質。其次,我們表明對比解碼透過利用同一家族模型中存在的類似偏差,有效地緩解了分數範圍偏差。揭示並解決這些偏差有助於 LLM 評審更好地與人類評估者對齊,並釋放擴展到標準 1-5 分數範圍之外的潛力。未來的研究方向包括擴展到超過 14B 參數的模型和摘要之外的任務,以及研究替代的輕量級方法,例如在評估量規上的提示工程 (Prompt Engineering)。
局限性
推論時間計算
對比學習由於需要在兩個模型上運行前向傳遞 (Forward Pass) 而非一個,增加了測試時間計算。另一方面,使用主模型和輔助模型在實際應用中非常常見,用於透過推測性解碼 (Speculative Decoding) 加速解碼 (Leviathan et al., 2023),因此在使用推測性解碼時,對比解碼可以不需要額外的前向傳遞。
模型大小
由於計算預算限制,我們的實驗僅限於最多 14B 參數的模型。
語言覆蓋
我們的實驗僅在英語上進行,不過我們並未利用英語特有的語言學知識。
任務覆蓋
我們的實驗遵循 Liu et al. (2023) 和 Panickssery et al. (2024) 在摘要任務上進行。然而,我們在摘要指標的多個維度上進行了實驗,即連貫性、相關性(附錄 C)和一致性(附錄 C),以確認分數範圍偏差不僅發生在某一特定維度上。
附錄 A 評審提示詞
我們使用 Liu et al. (2023) 實驗中的以下提示詞。
【譯註:提示詞為英文模板,包含三個評估維度(連貫性、相關性、一致性)的評估指令。每個提示詞使用 {min_range} 和 {max_range} 模板變數來適應不同的分數範圍。】
附錄 B 超參數
我們對對比解碼的兩個超參數進行網格搜索 (Grid Search):1) 溫度 $t$ 和 2) 縮放常數 $\lambda$,搜索範圍如下:
- $\lambda = [0.01, 0.1, 0.5, 1.0]$
- $t = [0.5, 1.0, 2.0, 3.0, 4.0, 5.0]$
表 3:連貫性評估中各主模型和輔助模型對的對比解碼超參數設定
| 主模型 | 輔助模型 | 範圍 | $\lambda$ | $t$ |
|---|---|---|---|---|
| Llama 3.1 8B | Llama 3.2 3B | 0-4 | 0.01 | 1.0 |
| 1-5 | 1.0 | 0.5 | ||
| 2-6 | 1.0 | 0.5 | ||
| 3-7 | 0.01 | 5.0 | ||
| Llama 3.2 1B | 0-4 | 0.01 | 0.5 | |
| 1-5 | 0.1 | 5.0 | ||
| 2-6 | 0.1 | 2.0 | ||
| 3-7 | 0.1 | 2.0 | ||
| Qwen 2.5 7B | Qwen 2.5 3B | 0-4 | 0.1 | 4.0 |
| 1-5 | 0.01 | 5.0 | ||
| 2-6 | 0.1 | 4.0 | ||
| 3-7 | 0.01 | 0.5 | ||
| Qwen 2.5 14B | Qwen 2.5 3B | 0-4 | 0.1 | 2.0 |
| 1-5 | 0.01 | 4.0 | ||
| 2-6 | 0.1 | 1.0 | ||
| 3-7 | 0.1 | 2.0 |
表 4:相關性評估中 Llama-3 各主模型和輔助模型對的超參數設定
| 主模型 | 輔助模型 | 範圍 | $\lambda$ | $t$ |
|---|---|---|---|---|
| Llama 3.1 8B | Llama 3.2 3B | 0-4 | 0.01 | 0.5 |
| 1-5 | 0.01 | 0.5 | ||
| 2-6 | 0.01 | 0.5 | ||
| 3-7 | 0.5 | 0.5 | ||
| Llama 3.2 1B | 0-4 | 0.01 | 0.5 | |
| 1-5 | 0.1 | 5.0 | ||
| 2-6 | 0.1 | 5.0 | ||
| 3-7 | 0.01 | 0.5 | ||
| Qwen 2.5 7B | Qwen 2.5 3B | 0-4 | 0.1 | 5.0 |
| 1-5 | 0.5 | 1.0 | ||
| 2-6 | 1.0 | 1.0 | ||
| 3-7 | 0.1 | 4.0 | ||
| Qwen 2.5 14B | Qwen 2.5 3B | 0-4 | 0.01 | 3.0 |
| 1-5 | 0.1 | 0.5 | ||
| 2-6 | 0.01 | 3.0 | ||
| 3-7 | 0.01 | 3.0 |
表 5:一致性評估中 Llama-3 各主模型和輔助模型對的超參數設定
| 主模型 | 輔助模型 | 範圍 | $\lambda$ | $t$ |
|---|---|---|---|---|
| Llama 3.1 8B | Llama 3.2 3B | 0-4 | 0.01 | 5.0 |
| 1-5 | 0.1 | 2.0 | ||
| 2-6 | 0.1 | 2.0 | ||
| 3-7 | 0.1 | 3.0 | ||
| Llama 3.2 1B | 0-4 | 0.1 | 1.0 | |
| 1-5 | 0.1 | 0.5 | ||
| 2-6 | 0.1 | 2.0 | ||
| 3-7 | 0.1 | 1.0 | ||
| Qwen 2.5 7B | Qwen 2.5 3B | 0-4 | 0.1 | 3.0 |
| 1-5 | 0.01 | 5.0 | ||
| 2-6 | 0.01 | 2.0 | ||
| 3-7 | 0.01 | 5.0 | ||
| Qwen 2.5 14B | Qwen 2.5 3B | 0-4 | 0.1 | 1.0 |
| 1-5 | 0.1 | 5.0 | ||
| 2-6 | 0.1 | 3.0 | ||
| 3-7 | 0.1 | 3.0 |
附錄 C 相關性與一致性結果
表 6:Llama-3 家族在摘要相關性上與人類標註的相關性結果
| 模型 | 範圍 | Pear. | Spear. | Kend. |
|---|---|---|---|---|
| Llama 3.2-1B-Inst | 0 to 4 | 0.336 | 0.317 | 0.283 |
| Llama 3.2-3B-Inst | 0 to 4 | 0.113 | 0.110 | 0.096 |
| Llama 3.1-8B-Inst | 0 to 4 | 0.420 | 0.385 | 0.340 |
| Contrastive (8B-1B) | 0 to 4 | 0.406 | 0.371 | 0.326 |
| Contrastive (8B-3B) | 0 to 4 | 0.403 | 0.370 | 0.327 |
| Llama 3.2-1B-Inst | 1 to 5 | 0.071 | 0.063 | 0.056 |
| Llama 3.2-3B-Inst | 1 to 5 | 0.138 | 0.148 | 0.130 |
| Llama 3.1-8B-Inst | 1 to 5 | 0.407 | 0.372 | 0.327 |
| Contrastive (8B-1B) | 1 to 5 | 0.399 | 0.374 | 0.331 |
| Contrastive (8B-3B) | 1 to 5 | 0.393 | 0.356 | 0.315 |
| Llama 3.2-1B-Inst | 2 to 6 | 0.000 | 0.000 | 0.000 |
| Llama 3.2-3B-Inst | 2 to 6 | 0.104 | 0.091 | 0.080 |
| Llama 3.1-8B-Inst | 2 to 6 | 0.305 | 0.288 | 0.255 |
| Contrastive (8B-1B) | 2 to 6 | 0.420 | 0.388 | 0.340 |
| Contrastive (8B-3B) | 2 to 6 | 0.325 | 0.305 | 0.270 |
| Llama 3.2-1B-Inst | 3 to 7 | 0.069 | 0.032 | 0.030 |
| Llama 3.2-3B-Inst | 3 to 7 | 0.099 | 0.096 | 0.082 |
| Llama 3.1-8B-Inst | 3 to 7 | 0.386 | 0.372 | 0.325 |
| Contrastive (8B-1B) | 3 to 7 | 0.395 | 0.376 | 0.330 |
| Contrastive (8B-3B) | 3 to 7 | 0.429 | 0.396 | 0.340 |
| 所有分數範圍的平均值 | ||||
| Llama 3.2-1B-Inst | 0.119 | 0.103 | 0.092 | |
| Llama 3.2-3B-Inst | 0.113 | 0.111 | 0.097 | |
| Llama 3.1-8B-Inst | 0.379 | 0.354 | 0.312 | |
| Contrastive (8B-1B) | 0.405 | 0.377 | 0.332 | |
| Contrastive (8B-3B) | 0.388 | 0.357 | 0.313 |
表 7:Qwen2.5 家族在摘要相關性上與人類標註的相關性結果
| 模型 | 範圍 | Pear. | Spear. | Kend. |
|---|---|---|---|---|
| Qwen2.5-3B-Inst | 0 to 4 | 0.000 | 0.000 | 0.000 |
| Qwen2.5-7B-Inst | 0 to 4 | 0.304 | 0.282 | 0.247 |
| Qwen2.5-14B-Inst | 0 to 4 | 0.474 | 0.451 | 0.398 |
| Contrastive (7B-3B) | 0 to 4 | 0.329 | 0.284 | 0.246 |
| Contrastive (14B-3B) | 0 to 4 | 0.502 | 0.487 | 0.425 |
| Qwen2.5-3B-Inst | 1 to 5 | 0.000 | 0.000 | 0.000 |
| Qwen2.5-7B-Inst | 1 to 5 | 0.337 | 0.332 | 0.284 |
| Qwen2.5-14B-Inst | 1 to 5 | 0.489 | 0.467 | 0.411 |
| Contrastive (7B-3B) | 1 to 5 | 0.349 | 0.327 | 0.285 |
| Contrastive (14B-3B) | 1 to 5 | 0.469 | 0.448 | 0.394 |
| Qwen2.5-3B-Inst | 2 to 6 | 0.000 | 0.000 | 0.000 |
| Qwen2.5-7B-Inst | 2 to 6 | 0.232 | 0.222 | 0.195 |
| Qwen2.5-14B-Inst | 2 to 6 | 0.436 | 0.412 | 0.362 |
| Contrastive (7B-3B) | 2 to 6 | 0.247 | 0.261 | 0.225 |
| Contrastive (14B-3B) | 2 to 6 | 0.445 | 0.426 | 0.374 |
| Qwen2.5-3B-Inst | 3 to 7 | 0.000 | 0.000 | 0.000 |
| Qwen2.5-7B-Inst | 3 to 7 | 0.193 | 0.230 | 0.207 |
| Qwen2.5-14B-Inst | 3 to 7 | 0.448 | 0.419 | 0.366 |
| Contrastive (7B-3B) | 3 to 7 | 0.244 | 0.263 | 0.234 |
| Contrastive (14B-3B) | 3 to 7 | 0.483 | 0.469 | 0.407 |
| 所有分數範圍的平均值 | ||||
| Qwen2.5-3B-Inst | 0.000 | 0.000 | 0.000 | |
| Qwen2.5-7B-Inst | 0.267 | 0.267 | 0.233 | |
| Qwen2.5-14B-Inst | 0.462 | 0.437 | 0.384 | |
| Contrastive (7B-3B) | 0.292 | 0.284 | 0.248 | |
| Contrastive (14B-3B) | 0.475 | 0.458 | 0.400 |
表 8:Llama-3 家族在摘要一致性上與人類標註的相關性結果
| 模型 | 範圍 | Pear. | Spear. | Kend. |
|---|---|---|---|---|
| Llama 3.2-1B-Inst | 0 to 4 | -0.039 | -0.027 | -0.024 |
| Llama 3.2-3B-Inst | 0 to 4 | 0.118 | 0.115 | 0.110 |
| Llama 3.1-8B-Inst | 0 to 4 | 0.494 | 0.454 | 0.435 |
| Contrastive (8B-1B) | 0 to 4 | 0.482 | 0.435 | 0.418 |
| Contrastive (8B-3B) | 0 to 4 | 0.450 | 0.414 | 0.396 |
| Llama 3.2-1B-Inst | 1 to 5 | -0.152 | -0.229 | -0.215 |
| Llama 3.2-3B-Inst | 1 to 5 | 0.208 | 0.213 | 0.202 |
| Llama 3.1-8B-Inst | 1 to 5 | 0.675 | 0.581 | 0.560 |
| Contrastive (8B-1B) | 1 to 5 | 0.672 | 0.578 | 0.557 |
| Contrastive (8B-3B) | 1 to 5 | 0.656 | 0.577 | 0.557 |
| Llama 3.2-1B-Inst | 2 to 6 | 0.000 | 0.000 | 0.000 |
| Llama 3.2-3B-Inst | 2 to 6 | 0.171 | 0.157 | 0.151 |
| Llama 3.1-8B-Inst | 2 to 6 | 0.476 | 0.449 | 0.434 |
| Contrastive (8B-1B) | 2 to 6 | 0.531 | 0.479 | 0.463 |
| Contrastive (8B-3B) | 2 to 6 | 0.548 | 0.497 | 0.480 |
| Llama 3.2-1B-Inst | 3 to 7 | -0.000 | 0.007 | 0.008 |
| Llama 3.2-3B-Inst | 3 to 7 | 0.195 | 0.180 | 0.172 |
| Llama 3.1-8B-Inst | 3 to 7 | 0.555 | 0.509 | 0.483 |
| Contrastive (8B-1B) | 3 to 7 | 0.540 | 0.492 | 0.466 |
| Contrastive (8B-3B) | 3 to 7 | 0.550 | 0.510 | 0.488 |
| 所有分數範圍的平均值 | ||||
| Llama 3.2-1B-Inst | -0.048 | -0.062 | -0.058 | |
| Llama 3.2-3B-Inst | 0.173 | 0.166 | 0.159 | |
| Llama 3.1-8B-Inst | 0.550 | 0.498 | 0.478 | |
| Contrastive (8B-1B) | 0.557 | 0.496 | 0.476 | |
| Contrastive (8B-3B) | 0.551 | 0.500 | 0.480 |
表 9:Qwen2.5 家族在摘要一致性上與人類標註的相關性結果
| 模型 | 範圍 | Pear. | Spear. | Kend. |
|---|---|---|---|---|
| Qwen2.5-3B-Inst | 0 to 4 | -0.486 | -0.485 | -0.474 |
| Qwen2.5-7B-Inst | 0 to 4 | 0.402 | 0.405 | 0.377 |
| Qwen2.5-14B-Inst | 0 to 4 | 0.452 | 0.456 | 0.430 |
| Contrastive (7B-3B) | 0 to 4 | 0.460 | 0.415 | 0.383 |
| Contrastive (14B-3B) | 0 to 4 | 0.451 | 0.453 | 0.428 |
| Qwen2.5-3B-Inst | 1 to 5 | -0.282 | -0.214 | -0.208 |
| Qwen2.5-7B-Inst | 1 to 5 | 0.560 | 0.516 | 0.487 |
| Qwen2.5-14B-Inst | 1 to 5 | 0.482 | 0.479 | 0.459 |
| Contrastive (7B-3B) | 1 to 5 | 0.565 | 0.510 | 0.484 |
| Contrastive (14B-3B) | 1 to 5 | 0.564 | 0.539 | 0.521 |
| Qwen2.5-3B-Inst | 2 to 6 | -0.175 | -0.175 | -0.170 |
| Qwen2.5-7B-Inst | 2 to 6 | 0.516 | 0.462 | 0.441 |
| Qwen2.5-14B-Inst | 2 to 6 | 0.165 | 0.171 | 0.158 |
| Contrastive (7B-3B) | 2 to 6 | 0.498 | 0.470 | 0.448 |
| Contrastive (14B-3B) | 2 to 6 | 0.280 | 0.313 | 0.292 |
| Qwen2.5-3B-Inst | 3 to 7 | -0.456 | -0.444 | -0.432 |
| Qwen2.5-7B-Inst | 3 to 7 | 0.454 | 0.470 | 0.448 |
| Qwen2.5-14B-Inst | 3 to 7 | 0.260 | 0.244 | 0.233 |
| Contrastive (7B-3B) | 3 to 7 | 0.457 | 0.469 | 0.447 |
| Contrastive (14B-3B) | 3 to 7 | 0.343 | 0.306 | 0.289 |
| 所有分數範圍的平均值 | ||||
| Qwen2.5-3B-Inst | -0.350 | -0.329 | -0.321 | |
| Qwen2.5-7B-Inst | 0.483 | 0.463 | 0.438 | |
| Qwen2.5-14B-Inst | 0.340 | 0.338 | 0.320 | |
| Contrastive (7B-3B) | 0.495 | 0.466 | 0.441 | |
| Contrastive (14B-3B) | 0.408 | 0.403 | 0.382 |
附錄 D 模型大小與預算
本論文的所有實驗均使用 NVIDIA A100 GPU。本論文使用的基礎模型在以下許可下授權:Llama 3 家族模型使用 Meta Llama 3 License,Qwen 2.5 家族模型使用 Apache-2.0 許可。我們遵循了其預期用途。
附錄 E AI 輔助工具使用說明
我們在本手稿中使用了 Claude 來提升論文的清晰度和修正語法錯誤。我們也使用它來建立實驗程式碼。
附錄 F 潛在風險
如局限性部分所述,實驗僅在英語上進行,這可能使結論偏向英語。
參考文獻
Bubeck et al. (2023) Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. Preprint, arXiv:2303.12712.
Chiang and Lee (2023) Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15607–15631, Toronto, Canada. Association for Computational Linguistics.
Fabbri et al. (2021) Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
Gambardella et al. (2024) Andrew Gambardella, Yusuke Iwasawa, and Yutaka Matsuo. 2024. Language models do hard arithmetic tasks easily and hardly do easy arithmetic tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 85–91, Bangkok, Thailand. Association for Computational Linguistics.
Goel et al. (2025) Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, Ameya Prabhu, Matthias Bethge, and Jonas Geiping. 2025. Great models think alike and this undermines ai oversight. Preprint, arXiv:2502.04313.
Grattafiori et al. (2024) Aaron Grattafiori et al. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783.
Jones et al. (2024) Jaylen Jones, Lingbo Mo, Eric Fosler-Lussier, and Huan Sun. 2024. A multi-aspect framework for counter narrative evaluation using large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 147–168, Mexico City, Mexico. Association for Computational Linguistics.
Kim et al. (2024) Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. Prometheus 2: An open source language model specialized in evaluating other language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 4334–4353, Miami, Florida, USA. Association for Computational Linguistics.
Leviathan et al. (2023) Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org.
Li et al. (2024) Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, hai zhao, and Pengfei Liu. 2024. Generative judge for evaluating alignment. In The Twelfth International Conference on Learning Representations.
Li et al. (2023) Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive decoding: Open-ended text generation as optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12286–12312, Toronto, Canada. Association for Computational Linguistics.
Likert (1932) Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of psychology.
Lin et al. (2022) Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics.
Liu et al. (2023) Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Liu et al. (2024) Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vulić, Anna Korhonen, and Nigel Collier. 2024. Aligning with human judgement: The role of pairwise preference in large language model evaluators. In First Conference on Language Modeling.
Nogueira et al. (2021) Rodrigo Nogueira, Zhiying Jiang, and Jimmy Lin. 2021. Investigating the limitations of transformers with simple arithmetic tasks. Preprint, arXiv:2102.13019.
O'Brien and Lewis (2023) Sean O'Brien and Mike Lewis. 2023. Contrastive decoding improves reasoning in large language models. Preprint, arXiv:2309.09117.
Panickssery et al. (2024) Arjun Panickssery, Samuel R. Bowman, and Shi Feng. 2024. LLM evaluators recognize and favor their own generations. In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
Qwen et al. (2025) Qwen, An Yang, Baosong Yang, et al. 2025. Qwen2.5 technical report. Preprint, arXiv:2412.15115.
Wataoka et al. (2024) Koki Wataoka, Tsubasa Takahashi, and Ryokan Ri. 2024. Self-preference bias in LLM-as-a-judge. In Neurips Safe Generative AI Workshop 2024.
Ye et al. (2024) Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V Chawla, and Xiangliang Zhang. 2024. Justice or prejudice? quantifying biases in llm-as-a-judge. Preprint, arXiv:2410.02736.
Zheng et al. (2023) Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS '23, Red Hook, NY, USA. Curran Associates Inc.
Zhu et al. (2025) Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2025. JudgeLM: Fine-tuned large language models are scalable judges. In The Thirteenth International Conference on Learning Representations.
術語對照表
| 英文術語 | 中文翻譯 |
|---|---|
| Contrastive Decoding | 對比解碼 |
| Score Range Bias | 分數範圍偏差 |
| LLM-as-a-Judge | LLM 評審 |
| Direct Assessment | 直接評估 |
| Pairwise Comparison | 成對比較 |
| Pointwise Assessment | 逐點評估 |
| Self-Enhancement Bias | 自我增強偏差 |
| Family-Enhancement Bias | 家族增強偏差 |
| Main Model | 主模型 |
| Assistant Model | 輔助模型 |
| Greedy Decoding | 貪婪解碼 |
| Hyperparameter | 超參數 |
| Grid Search | 網格搜索 |
| Forward Pass | 前向傳遞 |
| Speculative Decoding | 推測性解碼 |
| Prompt Engineering | 提示工程 |
| Coherence | 連貫性 |
| Relevance | 相關性 |
| Consistency | 一致性 |
| Likert Scale | Likert 量表 |
| Clamp | 截斷 |
| Summarization | 摘要 |