學習即遺忘:LLM 訓練作為有損壓縮
原始論文:Learning is Forgetting: LLM Training As Lossy Compression 作者:Henry C. Conklin, Tom Hosking, Tan Yi-Chern, Julian Gold, Jonathan D. Cohen, Thomas L. Griffiths, Max Bartolo, Seraphina Goldfarb-Tarrant arXiv ID:2604.07569v1 日期:2026-04-08 標籤:LLM Information Theory Representation Learning Information Bottleneck Compression 可解釋性 預訓練動態
摘要
儘管大型語言模型 (LLM) 的使用日益普及,我們對其表徵空間 (representational space) 的結構仍然缺乏理解。這限制了我們詮釋 LLM 學到了什麼、如何學習,以及將其與人類學習做類比的能力。本文主張,LLM 最適合被理解為有損壓縮 (lossy compression) 的實例——在訓練過程中,模型只保留與其目標相關的訓練資料資訊。我們展示了預訓練的結果是模型被最佳化壓縮以進行下一個序列預測,趨近資訊瓶頸 (Information Bottleneck) 的理論界限。在一系列開放權重模型中,每個模型的壓縮方式不同,可能是因為資料和訓練方法的差異。然而,跨越不同模型家族,一個模型壓縮的最佳程度以及它保留的資訊,能夠預測其在六個基準測試上的下游效能。整體而言,本文為 LLM 的學習方式提供了一個統一的資訊理論框架,並且可以在規模化部署。
1. 引言
我們對大型語言模型如何在廣泛任務上取得驚人成果的理解仍然有限。雖然有大量研究透過行為實驗、探測 (probing) 或因果干預來解讀 LLM,但模型的規模使得理解其表徵空間的結構是一項持續的挑戰。本文將 LLM 視為有損壓縮的實例,提供一個關於模型在訓練期間如何表徵資訊以及什麼資訊對效能重要的說明。
有損壓縮透過只保留與目標相關的資訊來高效地表徵資料。未壓縮的音訊可以是 GB 等級,MP3 透過丟棄人耳聽不到的頻率來節省空間;JPEG 同樣省略人眼難以察覺的色彩變化。我們在 LLM 中看到了類似的平行關係:LLM 在數兆 token 上訓練後——比一個人在 200 輩子中能接觸到的語言資料還多——被期望產生人類偏好的回應。
壓縮本質上是「有立場的」——來源中的某些資訊被保留,某些被遺忘以節省空間。資訊理論 (Information Theory, Shannon, 1948) 提供了一個正式框架來描述這個過程,讓我們既能量化表徵中的資訊,也能計算它相對於所表徵資料的最佳壓縮界限。我們的結果建立在深度學習的資訊瓶頸 (IB) 理論 (Tishby & Zaslavsky, 2015) 之上,展示預訓練遵循兩階段軌跡:先增加與訓練目標的互資訊 (mutual information),然後壓縮輸入資訊。
本文聚焦三個核心問題:LLM 是否最佳化壓縮其表徵?壓縮後存留下來的是什麼資訊?什麼樣的表徵結構驅動效能?
核心發現如下:
- 預訓練動態密切遵循資訊瓶頸的理論預測,模型先擴展表徵,然後慢慢趨近最佳壓縮
- 規模很重要:小於 70 億參數的模型在訓練後期難以達成有意義的壓縮
- 模型壓縮的最佳程度與跨六個基準測試的效能顯著相關($r = 0.52$,$p < 0.001$)
- 表徵中的偏好資訊 (preference information) 量能顯著預測下游效能($r = 0.76$,$p < 0.001$)
- 來自 5 個模型家族的大量開放權重模型,全都收斂到最佳壓縮的附近
2. 背景與相關工作
2.1 學習、推論與壓縮
壓縮一直被認為是人類學習和推論的基礎 (Chater, 1997; Feldman, 2000; Chater & Vitányi, 2003)。越來越多人將概率論式推論和複雜度最小化視為密切相關——這一點在 Bayesian 推論中表現得尤為清楚:Bayesian 推論隱含地偏好與觀察資料一致的最簡單假設。在機器學習中,Occam's Razor 長期被用作模型選擇準則,偏差-方差權衡 (bias-variance trade-off) 使這一點在神經網路中更為明確:更複雜的模型可能更好地擬合訓練資料,但也更容易泛化不佳。
雖然有一些工作研究 LLM 是否能做到無損壓縮 (lossless compression)(即上下文中的壓縮演算法),本文探討的是不同的面向——將 LLM 訓練本身視為一個有損壓縮的過程。
2.2 率失真理論
考慮一個函數 $\theta$ 將輸入 $X$ 編碼為表徵 $Z = \theta(X)$,再由解碼函數 $\phi$ 預測輸出 $\hat{Y} = \phi(Z)$。率失真理論 (Rate Distortion Theory, Shannon, 1948) 考慮的是有損情況 $\hat{Y} \neq Y$,即允許一定程度的預測誤差(失真 (distortion)),問題變成:編碼器需要保留多少輸入資訊(稱為速率 (rate))才能達到給定的失真水準。
資訊瓶頸 (Information Bottleneck, IB)。Tishby et al. (2000) 從率失真理論的一個特殊情況出發,定義速率為輸入與表徵之間的互資訊 $I(X; Z)$,失真為表徵與目標預測之間的互資訊 $I(Y; Z)$。由此構成的二維空間稱為資訊平面 (information plane)(見 Figure 1)。$I(X; Z)$ 反映保留了多少輸入資訊,稱為複雜度 (complexity);$I(Y; Z)$ 反映表徵對預測輸出有多少資訊,稱為表達性 (expressivity)。
IB 的目標函數為:
$$F_\beta[p(Z|X)] = I(X; Z) - \beta I(Y; Z) \tag{1}$$
其中 $\beta$ 是控制允許失真程度的權衡參數。透過變化 $\beta$,可以追蹤出一條界限曲線:曲線上方(編碼 $p(Z|X)$ 是最佳壓縮的)不可達,曲線下方皆為次最佳。
兩階段預測。Tishby & Zaslavsky (2015) 以及 Shwartz-Ziv & Tishby (2017) 提出訓練多層神經網路可以被理解為最佳化 IB:訓練分為兩個階段——(1) 擬合階段 (fitting phase),表徵增加與目標的互資訊 $I(Y; Z)$;(2) 壓縮階段 (compression phase),模型壓縮關於輸入的無關資訊 $I(X; Z)$,趨近最佳界限。後者被認為產生能良好泛化的表徵。
2.3 解讀神經網路
關於深度學習的學習動態,有大量文獻研究小型多層網路中的學習動態。然而,LLM 上的可解釋性工作大多依賴行為式或探測式方法,或是機械式可解釋性 (mechanistic interpretability)——描述模型內部的電路如何實現特定功能。本文的方法不同:我們將模型作為一個整體來解讀,而非聚焦於個別電路、注意力頭或神經元,並將結論置於既有的學習和壓縮理論框架中。
3. 方法
3.1 熵估計
設 $T \in \mathbb{Z}^{B \times S}$ 為一批 $B$ 個 token 化樣本(序列長度 $S$),模型 $\theta$ 有 $L$ 層、表徵維度 $h$,對應的編碼表徵為 $Z \in \mathbb{R}^{L \times B \times S \times h}$。我們使用軟熵估計器 (soft-entropy estimator, Conklin, 2025) 來計算互資訊 $I(X; Z)$,這是一種基於 Shannon 熵(而非微分熵)的可微分鬆弛近似,能夠擴展到 LLM 規模。
具體步驟:
- 正規化:將每個表徵向量 $z$ 歸一化到單位球面 $\mathbb{S}^h$ 上:$\bar{z} = z / \|z\|$
- 取樣錨點:從 $\mathbb{S}^h$ 上均勻取樣 $n$ 個點 $\{w_i\}_{i=1}^n$
- 軟分配:計算每個歸一化表徵 $\bar{z}$ 與各錨點 $w_i$ 的餘弦相似度,再透過帶溫度參數 $\epsilon$ 的 softmax 將每個表徵軟分配給各錨點:
$$\hat{Z}_{l,b,s,:} = \text{softmax}\left(\frac{\sum_{j=1}^{S} \hat{Z}_{l,b,s,j} W_{j,:}}{\epsilon}\right) \tag{2}$$
- 聚合:在 batch 和序列維度上取平均,得到每層 $l$ 的概率向量 $\hat{z}_l$
- 估計熵:計算每層的 Shannon 熵 $H(\hat{z}_l) = -\sum_j \hat{z}_{l,j} \log \hat{z}_{l,j}$
整個模型的熵透過對各層取平均並除以均勻分布的熵 $\log(n)$ 來歸一化,得到效率 (efficiency) 指標 $\mathcal{H}(Z)$:
$$\mathcal{H}(Z) := \frac{1}{L \log(n)} \sum_{l=1}^{L} H(\hat{z}_l) \tag{4}$$
條件熵 $\mathcal{H}(Z|X = x)$ 的計算方式類似,但僅在對應輸入 $x$ 的表徵上估計。最終互資訊為:
$$I(X; Z) := \mathcal{H}(Z) - \sum_{x \in X} P(X = x) \, \mathcal{H}(Z | X = x) \tag{5}$$
[Figure 2: 軟熵估計的示意圖。(上) 展示歸一化、取樣錨點、軟分配三個步驟。(下) 軟分配結果被聚合成描述空間 $P(\hat{Z})$ 的分布,然後計算其 Shannon 熵。]
3.2 互資訊與回退
為了判斷模型是否相對於輸入和輸出標籤達到最佳壓縮,需要計算關於輸入和輸出的互資訊。LLM 以前文作為輸入、後續文本作為輸出來訓練。為每個可能的上下文視窗維護條件估計 $P(Z|X)$ 在 LLM 規模下不可行,因此借用語言模型中 n-gram 回退 (back-off) 的思路,用有限寬度的前文來近似 $P(Z|X)$。
具體來說,使用 token、bigram、trigram 和 quadgram(四元語法)等不同層級的上下文回退。隨著上下文增加,表徵趨近界限的程度也更高。
定義壓縮最佳性 (optimality) 為:
$$\text{Optimality} = \frac{\text{Expressivity}}{\text{Complexity}} = \frac{I(Y; Z)}{I(X; Z)} \tag{6}$$
當表徵系統趨近界限時,此值趨近 1.0。
[Figure 3: 條件概率估計的示意圖。以一個範例句子展示如何用 bigram 回退來估計條件互資訊。]
除了輸入/輸出的互資訊外,本文還考慮偏好資訊 (preference information)。使用偏好資料(一個提示有兩個延續,一個被標記為偏好、另一個被拒絕),條件化在偏好標籤上計算 $P(Z|\text{preferred})$ 和 $I(Z; \text{preferred})$。
資料與取樣:除非另有說明,token、bigram、trigram 和 quadgram 估計使用 C4 資料集的 10,000 個樣本,偏好估計使用 Tulu 資料集的 10,000 個樣本,最大上下文長度 512。
4. 實驗
4.1 預訓練趨近最佳壓縮
預訓練的大部分過程看起來是訓練資料的緩慢壓縮。資訊瓶頸的深度學習理論預測兩個階段:(1) 擬合階段,輸出互資訊 $I(Y; Z)$ 增加;(2) 壓縮階段,輸入互資訊 $I(X; Z)$ 降低,表徵趨近最佳界限。
[Figure 1 (左): OLMo2 7B 模型的資訊平面。橫軸為複雜度 $I(X; Z)$,縱軸為表達性 $I(Y; Z)$。虛線為最佳壓縮界限。色調表示訓練進度(以 token 數計,十億為單位)。模型先向右上擴展,再向左上壓縮,趨近界限。]
[Figure 1 (右): OLMo2 7B 的交叉熵損失與壓縮最佳性的關係。隨著損失趨於飽和,表徵開始趨近界限。]
以 OLMo2 7B 為例,模型密切遵循 IB 的兩階段預測:先增加與輸出的互資訊,然後壓縮輸入資訊,逐步趨近最佳壓縮界限。而且,這個轉折發生在模型的下一個 token 預測損失開始飽和時。
模型主要編碼局部上下文。透過分析不同回退層級,發現模型中大部分資訊編碼的是局部上下文(token 到 quadgram 層級)。1B 模型比更大的模型有更多的 token 層級資訊和更少的上下文資訊。
[Figure 4: (上) 不同回退層級的資訊平面。所有上下文寬度都呈現相同的兩階段模式,但更多上下文的表徵更趨近界限。(下) 各模型規模(1B、7B、32B)中不同回退層級所佔的資訊比例。]
規模的影響:小模型難以壓縮。參數量對模型能達到的壓縮程度有顯著影響。1B、7B 和 32B 模型在預訓練中呈現明顯不同的行為:
- 7B 和 32B 模型密切遵循 IB 軌跡,展現擴展和壓縮兩階段,最終趨近最佳壓縮
- 1B 模型成功完成擴展階段(增加 $I(Y; Z)$),但在壓縮階段失敗——它在界限附近震盪,無法趨近最佳壓縮
[Figure 5: (左上) 6 個家族的開放權重模型在訓練結束時都在界限附近。色調表示 MMLU Pro 分數,越綠越好。(右上) 偏好資訊 vs. 複雜度。偏好資訊越多的模型效能越好。(下) 1B、7B、32B 三種規模的預訓練軌跡。1B 模型在第二階段震盪,無法穩定壓縮。]
這暗示,要達到最佳壓縮,可能存在一個參數量的門檻——這與縮放定律 (scaling laws) 的觀察一致。
大量開放權重模型收斂到界限附近。在 OLMo2 家族之外,計算了來自多個不同家族的開放權重 LLM 的複雜度和表達性,發現它們都收斂到界限附近(Figure 5 左上,以及附錄 Figure 8、9 的 75 個模型)。這表明,趨近最佳壓縮不是某個特定訓練方法的產物,而是深度學習模型的基本屬性。
4.2 將表徵結構與效能連結
接下來考慮壓縮結構如何與下游效能相關。分析了來自 6 個家族的 47 個開放權重模型,跨越六個基準測試(MMLU Pro, BBH, Math LVL5, IFEval, GPQA, MuSR)。
[Figure 6: 表徵資訊與效能的關係。(左上) 較低的複雜度(token 回退)與較好的下游效能顯著相關($r = -0.38$,$p = 0.006$)。(右上) 表達性本身與效能的關係不顯著($r = 0.08$,$p = 0.575$)。(左下) 壓縮最佳性與效能顯著相關($r = 0.52$,$p < 0.001$)。(右下) 偏好資訊與效能的相關性最強($r = 0.76$,$p < 0.001$)。]
關鍵發現:
- 複雜度與效能負相關:token 層級的複雜度越低(壓縮越多),效能越好
- 表達性本身不預測效能:光有高表達性不夠
- 壓縮最佳性預測效能($r = 0.52$,$p < 0.001$):表達性與複雜度的比值越接近 1,效能越好
- 偏好資訊最能預測效能($r = 0.76$,$p < 0.001$):表徵中包含的偏好資訊量是下游效能的最強預測因子
更多上下文資訊 → 更好效能。在 token 層級之外,進一步分析 bigram、trigram、quadgram 各層級的條件互資訊。結果(Figure 7)顯示:
- 更少的 token 資訊與更好的效能相關
- 更多的 bigram、trigram、quadgram 資訊與更好的效能正相關
[Figure 7: 不同回退層級的資訊比例 vs. 效能。token 層級資訊與效能負相關,但 bigram、trigram、quadgram 等上下文資訊與效能正相關。更高的壓縮最佳性在所有層級都與更好效能正相關。]
也就是說,效能更好的模型把更少的表徵容量分配給 token 層級的區分,而把更多容量用於編碼更豐富的上下文資訊。
偏好資訊。雖然 LLM 在預訓練中趨近下一個 token 預測的最佳壓縮,但大量工作也在改善模型遵循指令和產生人類偏好回應的能力。使用偏好資料計算的偏好互資訊量,是下游效能的最強預測因子($r = 0.76$,$p < 0.001$)。附錄 B 進一步展示,後訓練 (post-training) 可以增加偏好資訊量而不顯著改變複雜度,表明預訓練負責廣泛的壓縮,而後訓練編輯的是壓縮所保留的具體資訊內容。
應用前景。這些結果暗示資訊理論方法可以在訓練中被利用:壓縮最佳性可作為停止準則(當距離界限不再減小時停止訓練),或作為模型選擇準則(選擇壓縮最佳或偏好資訊最多的 checkpoint)。這些估計基於單次前向傳播和教師強制 (teacher forcing) 計算,成本遠低於在一整套基準測試上評估模型。
5. 結論
本文彌合了學習的理論說明與 LLM 實際複雜性之間的差距。我們展示了 LLM 學習的是對訓練資料的最佳壓縮,大量開放權重模型沿著 IB 界限收斂——壓縮最佳性能預測下游效能。每個模型的壓縮方式不同,但我們能說明壓縮過程中存留了什麼資訊:表徵編碼了不同層級的局部上下文和人類偏好的資訊。
本文引入的可解釋性方法將模型作為整體來解讀——而非聚焦於特定電路、注意力頭或最終層的嵌入——因為複雜的分散式系統不能僅從其部件來理解。我們主張 LLM 最適合被理解為有損壓縮,並將其置於表徵學習這一跨科學領域的長期研究傳統中。
附錄摘要
- 附錄 A:75 個開放權重模型的 token 和 bigram 資訊平面完整視覺化(Figure 8),模型普遍趨近 IB 界限
- 附錄 B:後訓練對壓縮的影響——後訓練增加偏好資訊但不顯著改變複雜度
- 附錄 C:各基準測試的個別分析——壓縮最佳性顯著預測數學、推理和事實知識任務的效能,但不預測指令遵循;指令遵循由偏好資訊預測
- 附錄 D:SmolLM2 和 Pythia 模型家族的預訓練分析,呈現與 OLMo2 類似的模式
- 附錄 E:方法論細節(熵估計的校準、回退程序、標籤程序、條件互資訊的估計)
- 附錄 F:實驗中使用的所有模型完整列表
- 附錄 G:使用的計算資源
- 附錄 H:資料集與可重現性細節
參考文獻
Aggarwal, C. C., Hinneburg, A., & Keim, D. A. (2001). On the surprising behavior of distance metrics in high dimensional space. International Conference on Database Theory (ICDT), 420–434.
Allal, L. B., Lozhkov, A., Bakouch, E., et al. (2025). Smollm2: When smol goes big – data-centric training of a small language model. Mathematics of computation, 28(125), 239–251.
Anderson, P. W. (1972). More is different: Broken symmetry and the nature of the hierarchical structure of science. Science, 177(4047), 393–396.
Bai, Y., Jones, A., Ndousse, K., et al. (2022). Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862.
Biderman, S., Schoelkopf, H., Anthony, Q. G., et al. (2023). Pythia: A suite for analyzing large language models across training and scaling. International Conference on Machine Learning, 2397–2430.
Blahut, R. (1972). Computation of channel capacity and rate-distortion functions. IEEE transactions on Information Theory, 18(4), 460–473.
Bricken, T., Templeton, A., Batson, J., et al. (2023). Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2.
Burnham, K. P., & Anderson, D. R. (2002). Model selection and multimodel inference: A practical information-theoretic approach. Springer.
Chater, N. (1997). Simplicity and the mind. The Psychologist.
Chater, N., & Vitányi, P. (2003). Simplicity: A unifying principle in cognitive science? Trends in Cognitive Sciences, 7(1), 19–22.
Coleman, C., Yeh, C., Mussmann, S., et al. (2020). Selection via proxy: Efficient data selection for deep learning. In ICLR.
Conklin, H. C. (2025). Information structure in mappings: An approach to learning, representation and generalisation. The University of Edinburgh.
Délétang, G., Ruoss, A., Duquenne, P.-A., et al. (2023). Language modeling is compression. arXiv preprint arXiv:2309.10668.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 4171–4186.
Edwards, A. W. F. (1972). Likelihood. Springer.
Elhage, N., Hume, T., Olsson, C., et al. (2022). Toy models of superposition. arXiv preprint arXiv:2209.10652.
Elhage, N., Nanda, N., Olsson, C., et al. (2021). A mathematical framework for transformer circuits. Transformer Circuits Thread, 1, 1.
Feldman, J. (2000). Minimization of boolean complexity in human concept learning. Nature, 407(6804), 630–633.
Feldman, J. (2016). The simplicity principle in perception and cognition. Wiley Interdisciplinary Reviews: Cognitive Science, 7(5), 330–340.
Fourrier, C., Habib, N., Lozovskaya, A., Szafer, K., & Wolf, T. (2024). Open llm leaderboard v2.
Frankle, J., & Carbin, M. (2018). The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv preprint arXiv:1803.03635.
Futrell, R., Wilcox, E., Morita, T., & Levy, R. (2018). Rnns as psycholinguistic subjects: Syntactic state and grammatical dependency. arXiv preprint arXiv:1809.01329.
Futrell, R., Wilcox, E., Morita, T., Qian, P., Ballesteros, M., & Levy, R. (2019). Neural language models as psycholinguistic subjects: Representations of syntactic state. arXiv preprint arXiv:1903.03260.
Gao, L., Biderman, S., Black, S., et al. (2020). The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027.
Gao, T., Yao, X., & Chen, D. (2021). Simcse: Simple contrastive learning of sentence embeddings. In EMNLP, 6894–6910.
Ge, X., Shu, W., Wu, J., Zhou, Y., He, Z., & Qiu, X. (2025). Evolution of concepts in language model pre-training.
Geman, S., Bienenstock, E., & Doursat, R. (1992). Neural networks and the bias/variance dilemma. Neural computation, 4(1), 1–58.
Gibson, E. (1998). Linguistic complexity: Locality of syntactic dependencies. Cognition, 68(1), 1–76.
Gibson, E., et al. (2000). The dependency locality theory: A distance-based theory of linguistic complexity. Image, language, brain, 2000, 95–126.
Goldfeld, Z., Berg, E. v. d., Greenewald, K., et al. (2019). Estimating Information Flow in Deep Neural Networks. arXiv:1810.05728.
Grattafiori, A., Dubey, A., Jauhri, A., et al. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Griffiths, T. L., Chater, N., & Tenenbaum, J. B. (2024). Bayesian models of cognition: Reverse engineering the mind. MIT Press.
Hahn, M., Futrell, R., Levy, R., & Gibson, E. (2022). A resource-rational model of human processing of recursive linguistic structure. Proceedings of the National Academy of Sciences, 119(43), e2122602119.
Hendrycks, D., Burns, C., Basart, S., et al. (2020). Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300.
Hendrycks, D., Burns, C., Kadavath, S., et al. (2021). Measuring mathematical problem solving with the MATH dataset. NeurIPS.
Hu, J., Gauthier, J., Qian, P., Wilcox, E., & Levy, R. P. (2020). A Systematic Assessment of Syntactic Generalization in Neural Language Models. arXiv:2005.03692.
Jayant, N., Johnston, J., & Safranek, R. (1993). Signal compression based on models of human perception. Proceedings of the IEEE, 81(10), 1385–1422.
Jaynes, E. T. (1957). Information theory and statistical mechanics. Physical Review, 106, 620–630.
Jeffreys, H. (1939). The theory of probability. OuP Oxford.
Kaplan, J., McCandlish, S., Henighan, T., et al. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361.
Katz, S. (1987). Estimation of probabilities from sparse data for the language model component of a speech recognizer. IEEE transactions on acoustics, speech, and signal processing, 35(3), 400–401.
Kirby, S., Tamariz, M., Cornish, H., & Smith, K. (2015). Compression and communication in the cultural evolution of linguistic structure. Cognition, 141, 87–102.
Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tulu 3: Pushing frontiers in open language model post-training. Cambridge university press.
MacKay, D. J. (2003). Information theory, inference and learning algorithms. Cambridge university press.
Marvin, R., & Linzen, T. (2018). Targeted Syntactic Evaluation of Language Models. arXiv:1808.09031.
Mitchell, M. (2009). Complexity: A guided tour. Oxford University Press.
Nanda, N., Chan, L., Lieberum, T., Smith, J., & Steinhardt, J. (2023). Progress Measures for Grokking via Mechanistic Interpretability.
OLMo, T., Walsh, P., Soldaini, L., et al. (2025, January). 2 OLMo 2 Furious. arXiv:2501.00656.
Ouyang, L., Wu, J., Jiang, X., et al. (2022). Training language models to follow instructions with human feedback. In NeurIPS, pp. 27730–27744.
Oyama, M., Yokoi, S., & Shimodaira, H. (2023). Norm of word embedding encodes information gain. In EMNLP, 2108–2130.
Paninski, L. (2003). Estimation of Entropy and Mutual Information. Neural Computation, 15(6), 1191–1253.
Pimentel, T., Valvoda, J., Maudslay, R. H., et al. (2020). Information-theoretic probing for linguistic structure. arXiv preprint arXiv:2004.03061.
Poggio, T., Rifkin, R., Mukherjee, S., & Niyogi, P. (2004). General conditions for predictivity in learning theory. Nature, 428(6981), 419–422.
Pothos, E. M., & Chater, N. (2001). 4 categorization by simplicity: A minimum description length approach to unsupervised clustering.
Rafailov, R., Sharma, A., Mitchell, E., et al. (2023). Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, pp. 53728–53741.
Raffel, C., Shazeer, N., Roberts, A., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(140), 1–67.
Reimers, N., & Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. In EMNLP, 3982–3992.
Rein, D., Hou, B. L., Stickland, A. C., et al. (2024). GPQA: A graduate-level google-proof q&a benchmark. COLM.
Rissanen, J. (1978). Modeling by shortest data description. Automatica, 14(5), 465–471.
Sajjadi, M. S., Bachem, O., Lucic, M., Bousquet, O., & Gelly, S. (2018). Assessing generative models via precision and recall. In NeurIPS, 31.
Saxe, A. M., Bansal, Y., Dapello, J., et al. (2019). On the information bottleneck theory of deep learning. JSTAT, 2019(12), 124020.
Saxe, A. M., McClelland, J. L., & Ganguli, S. (2019). A mathematical theory of semantic development in deep neural networks. PNAS, 116(23), 11537–11546.
Shannon, C. E. (1948). A mathematical theory of communication. The Bell system technical journal, 27(3), 379–423.
Shwartz-Ziv, R., & Tishby, N. (2017). Opening the black box of deep neural networks via information. arXiv preprint arXiv:1703.00810.
Sprague, Z., Ye, X., Bostrom, K., Chaudhuri, S., & Durrett, G. (2024). MuSR: Testing the limits of chain-of-thought with multistep soft reasoning. ICLR.
Suzgun, M., Scales, N., Schärli, N., et al. (2022). Challenging BIG-Bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
Tishby, N., Pereira, F. C., & Bialek, W. (2000). The information bottleneck method. arXiv preprint physics/0004057.
Tishby, N., & Zaslavsky, N. (2015). Deep learning and the information bottleneck principle. In IEEE information theory workshop (itw), 1–5.
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. arXiv preprint arXiv:1706.03762.
Veldhoen, S., Hupkes, D., & Zuidema, D. (2016). Diagnostic classifiers: Revealing how neural networks process hierarchical structure, 10.
Vitányi, P. M., & Li, M. (2000). Minimum description length induction, bayesianism, and kolmogorov complexity. IEEE Transactions on information theory, 46(2), 446–464.
Voita, E., Sennrich, R., & Titov, I. (2019). The Bottom-up Evolution of Representations in the Transformer. arXiv:1909.01380.
Voita, E., & Titov, I. (2020). Information-Theoretic Probing with Minimum Description Length. arXiv:2003.12298.
Wallace, C. S., & Boulton, D. M. (1968). An information measure for classification. The Computer Journal, 11(2), 185–194.
Wang, Y., Ma, X., Zhang, G., et al. (2024). Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In NeurIPS Datasets and Benchmarks Track.
Warstadt, A., Parrish, A., Liu, H., et al. (2019). BLiMP: A Benchmark of Linguistic Minimal Pairs for English. arXiv:1912.00582.
Wilcox, A. R. (1967). Indices of qualitative variation. Oak Ridge National Lab (ORNL).
Zaslavsky, N., Kemp, C., Regier, T., & Tishby, N. (2018). Efficient compression in color naming and its evolution. PNAS, 115(31), 7937–7942.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). Bertscore: Evaluating text generation with bert. In ICLR.
Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., & Hou, L. (2023). Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911.
術語對照表
| 英文 | 中文 |
|---|---|
| Lossy Compression | 有損壓縮 |
| Lossless Compression | 無損壓縮 |
| Information Bottleneck (IB) | 資訊瓶頸 |
| Rate Distortion Theory | 率失真理論 |
| Information Plane | 資訊平面 |
| Mutual Information | 互資訊 |
| Shannon Entropy | Shannon 熵 |
| Complexity | 複雜度 |
| Expressivity | 表達性 |
| Optimality | 最佳性 |
| Compression | 壓縮 |
| Distortion | 失真 |
| Rate | 速率 |
| Fitting Phase | 擬合階段 |
| Compression Phase | 壓縮階段 |
| Representation / Representational Space | 表徵 / 表徵空間 |
| Back-off | 回退 |
| N-gram | N 元語法 |
| Bigram / Trigram / Quadgram | 二元語法 / 三元語法 / 四元語法 |
| Soft-entropy Estimator | 軟熵估計器 |
| Preference Information | 偏好資訊 |
| Downstream Performance | 下游效能 |
| Open-weights Model | 開放權重模型 |
| Probing | 探測 |
| Mechanistic Interpretability | 機械式可解釋性 |
| Teacher Forcing | 教師強制 |
| Scaling Laws | 縮放定律 |
| Bias-Variance Trade-off | 偏差-方差權衡 |
| Occam's Razor | 奧卡姆剃刀 |
| Post-training | 後訓練 |
| Checkpoint | 檢查點 |
| Softmax | Softmax(歸一化指數函數) |
| Cosine Similarity | 餘弦相似度 |
| Anchor Points | 錨點 |
| Conditional Entropy | 條件熵 |
| Spearman Correlation | Spearman 等級相關 |