Jina-OCR-v1:以推測式解碼與稠密可驗證獎勵做高效文件解析
原始論文:Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards 作者:Alejandro Barón García, Feng Wang, Emilia Garcia Casademont, Han Xiao arXiv ID:2609.03181v1 日期:2026-09-02 標籤:OCR 文件解析 推測式解碼 MTP GRPO RLVR MoE 授權:CC BY-NC-SA 4.0;本譯文為原文的衍生作品,依相同條款釋出
摘要
我們提出 Jina-OCR-v1,一個為在低預算 GPU 上服務而打造的端到端文件解析模型。它結合 DeepSeek-OCR 的壓縮視覺編碼器 (Compressed-Vision Encoder) 與 3B 的 Mixture-of-Experts (MoE) decoder(每 token 約激活 570M 參數),再加上一個 FastMTP 推測式解碼 (Speculative Decoding) 頭——後者把單一 draft block 在 $K{=}3$ 個預測步之間遞迴共享。以貪婪驗證 (Greedy Verification) 讓解碼無損 (lossless)。後訓練結合指令對齊、對困難文件的穩健性微調,以及在稠密可驗證獎勵 (Dense Verifiable Rewards) 下的 GRPO——即對公式、表格與結構做確定性檢查、給予部分分數 (Partial Credit)。訓練資料混合清洗過的公開語料與有針對性的合成頁面。在預設的動態解析度設定下,Jina-OCR-v1 在 OmniDocBench v1.6 得 91.14 分、olmOCR-Bench 得 83.4 分,並在我們的比較中達到最高頁面吞吐 2.57 頁/秒。在 NVIDIA L4 這類低預算 GPU 上,FastMTP 相對貪婪自迴歸解碼把速度翻倍。模型公開於 https://huggingface.co/jinaai/jina-ocr-v1 。
1 引言
用視覺語言模型 (VLM) 做端到端文件解析近來成為一個重要範式,把版面分析、文字辨識、公式轉錄與表格結構還原摺疊進單一生成式流程。其輸出提供可供檢索、grounding 與代理式工作流使用的機器可讀表示,但生產使用——尤其在部署常見的低預算 GPU 上——仍受解碼成本與後訓練資料覆蓋所限。
長輸出讓解碼昂貴。 Qwen2.5-VL-72B 這類通用 VLM 每頁用數千個視覺 token 且嚴格自迴歸解碼。DeepSeek-OCR 證明壓縮視覺編碼器與精簡的 MoE decoder 能大幅降低服務成本。我們在該架構之上,瞄準剩下的自迴歸瓶頸:OCR 輸出的局部結構讓我們得以在驗證器評估之前先草擬 (draft) 好幾個 token。
後訓練需要可靠的監督與覆蓋。 開放標籤可能含退化迴圈 (degeneration loop) 或畸形結構;而公式與表格專屬的獎勵只適用於含這些元素的頁面。同時,文件請求橫跨整頁解析、元素轉錄與不同格式慣例。
我們的主要設計選擇是:
- FastMTP 解碼:把一個遞迴共享的 FastMTP draft block($K{=}3$)接到 DeepSeek-OCR 的壓縮視覺、3B-MoE 骨幹上。跨預測深度共享同一個 block,讓 draft 參數量在 $K$ 上保持常數。
- 稠密可驗證獎勵下的後訓練:以 ReMax baseline 對內容、結構、單元測試、重複與格式做乘法式 GRPO 獎勵最佳化。每一項都是「對照參考的確定性程式碼」並給予分級評分,因此部分正確的頁面能拿到部分分數。塞滿可評分公式與表格的合成頁面(JinaOCRSynth)提高了「公式與表格獎勵項可適用」的訓練頁面比例。
- 面向指令的訓練:在 ReaderLM-v2 之上,混合涵蓋多種解析風格、元素級轉錄、captioning、文件 VQA 與關鍵資訊抽取,並動態生成指令、附輔助 grounding 範例。

圖 1:特化 OCR 模型總覽(藍色標記本工作)。(a) 每視覺 token 的像素數(對數軸);(b) olmOCR-Bench 上的頁面吞吐(頁/秒,單張 A100,並發 32);(c) 效能對激活參數(對數軸),實線連接 Pareto 最佳系統。圖 1 把 Jina-OCR-v1 放在共同決定部署成本的三個軸上:視覺壓縮、頁面吞吐、每激活參數的準確度。
- (a) 視覺壓縮:有效 patch 邊長 $p$ 是一個視覺 token 的空間範圍,「每視覺 token 像素數」即該 token 代表多少影像像素。邊長 $p$ 的方形 patch ViT 約產出 $p^2$ 像素/token,故 Qwen-VL 家族用的 28–32 px 編碼器落在約 780–1,000 像素/token。DeepEncoder 則把 $1024{\times}1024$ 全域視圖的 4,096 個 patch 映射到 256 個 token($p{=}64$,約 4,096 像素/token)。Jina-OCR-v1 繼承該編碼器,故相近解析度的一頁用遠少的視覺 token 表示。
- (b) 頁面吞吐:單靠視覺壓縮不足以決定它——DeepSeek-OCR 的每 token 像素密度與 Jina-OCR-v1 相當,卻輸出更長的頁面。
- (c) 解析品質:Jina-OCR-v1 每 token 激活 570M decoder 參數,在兩個基準上都位於「該預算下的準確度前沿」,領先大它一個數量級的系統;在 OmniDocBench 上更是該前沿上最精簡的點。
2 相關工作
端到端文件解析。 Nougat 首先展示端到端學術文件 OCR,GOT-OCR2.0 把範圍擴到廣泛的 OCR-2.0 任務套件。近期系統包括 dots.ocr、SmolDocling、olmOCR、MonkeyOCR/MonkeyOCRv2、MinerU、PaddleOCR-VL、GLM-OCR、Infinity-Parser,以及本工作的架構基礎 DeepSeek-OCR。這份清單同時含端到端系統與「先跑版面模型(如 PP-DocLayout)」的兩階段管線。
高效解碼與多 token 預測。 推測式解碼平行驗證草擬的 token。從平行頭預測多個 token 的想法可追溯到 blockwise parallel decoding;Medusa 把它復活為凍結骨幹上的一組獨立 draft 頭,Hydra 讓這些頭依序相依。Lookahead decoding 則完全移除 draft 模型。多 token 預測 (MTP) 由 Gloeckle et al. 與 DeepSeek-V3 推廣為訓練目標;EAGLE 式方法在特徵層草擬。FastMTP 跨預測深度共享一個 draft block。GLM-OCR 已把共享參數 MTP 用於 OCR,HunyuanOCR-1.5 採用 DFlash block drafting。我們沿共享 block 路線(因其 draft 參數量恆定),並以特徵與分布對齊損失訓練它。
文件解析的強化學習。 GRPO 與 DAPO 已被調適到「可驗證獎勵強化學習 (RLVR)」下的文件解析,使用編輯距離、TEDS、CDM 等訊號。文件元素的結構化內容天生適合 RLVR,因為參考轉錄允許確定性評分、不需獎勵模型或裁判。RL 也比監督微調更不易災難性遺忘,這支持「用不同獎勵在子任務上特化」。我們結合乘法式獎勵加每項下限、ReMax baseline、顯式重複懲罰,以及集中在公式與表格結構的合成頁面。
3 模型架構
Jina-OCR-v1 沿用 DeepSeek-OCR 的 encoder–decoder 架構,並以一個多 token 預測頭擴展。
表 1:Jina-OCR-v1 模型規格。(DeepEncoder 與 MoE decoder 繼承自 DeepSeek-OCR;FastMTP 為新增)
| 組件 | 規格 |
|---|---|
| 視覺編碼器 | DeepEncoder (~380M):SAM (80M) → 16× conv → CLIP-L (300M) |
| 視覺 token | 256 @ 1024×1024 (Base);256+100n, n≤9 (Gundam,≤1,156/頁) |
| Decoder | DeepSeek-3B-MoE:12 層、$d{=}1280$、64 routed + 2 shared、top-6 |
| 激活/總參數 | ~570M / ~3B (decoder);<1B / ~3.4B (整個模型) |
| 詞彙量 | 129,280 |
| 位置上限 | 32,768(RoPE,$\theta{=}10^6$) |
| MTP 頭 | 1 個共享 dense block,遞迴 $K{=}3$ 步(FastMTP) |
DeepEncoder 把一個 window-attention 的 SAM 編碼器與一個 global-attention 的 CLIP 編碼器,透過 16× 卷積壓縮器串接。在 $1024{\times}1024$ 下它把 4,096 個全域 patch 壓到 256 個 token;動態解析度的 Gundam 模式最多再加九個 100-token 的局部 tile。Decoder 是一個精簡的 DeepSeekMoE 模型,每 token 激活約 570M 參數、輸出 Markdown。
圖 2(架構,文字描述):DeepEncoder 與 MoE decoder 沿用 DeepSeek-OCR——一頁產出 $1024{\times}1024$ 全域視圖的 256 個視覺 token,加上 $n$ 個各 100 token 的局部 tile。FastMTP 頭(橘色)從一個共享 draft block 遞迴提出 $K{=}3$ 個 token 供驗證器評估。
3.1 Draft 頭架構
OCR 輸出近乎確定、局部結構強,是有利的推測式解碼工作負載。我們採 FastMTP:單一 draft block $B_\theta$ 遞迴套用 $K{=}3$ 步,故 draft 參數量在 $K$ 上恆定。
訓練在「序列位置 × 深度」的二維格上對齊各遞迴深度。令 $x_p$ 為位置 $p$ 的 token、$e_p=E(x_p)$ 其嵌入、$b_p$ 為主 decoder 消化 $x_p$ 後的 post-norm 狀態,$u_p^{(k)}$ 為訓練列 $p$、遞迴深度 $k$ 的 draft 輸出。呈給共享頭的狀態經由 $N_e, N_h$ 正規化層、線性投影 $P$ 與共享 dense draft block $B_\theta$ 構成(釋出檢查點中即 enorm、hnorm、eh_proj、mtp_block)。token 嵌入 $E$、輸出 norm $N_o$ 與 LM head $W_{\mathrm{lm}}$ 綁定到固定的驗證器對應物。嵌入與 RoPE 位置跨深度固定,而預測狀態從第 $p-1$ 列移到第 $p$ 列。
每個深度,訓練列 $p$ 估計同一個驗證器狀態 $b_{p+1}$、其 logits 預測 $x_{p+2}$。看似的多 token 前進是對角的:對從第 $j$ 列起的 draft 鏈,深度 $k$ 佔 $p=j+k-1$,故它預測 $x_{j+k+1}$。
推論時我們遵循 EAGLE 的特徵級草擬範式,draft 頭從隱藏狀態(以及 token id)提出 token,推測式解碼迴圈用貪婪驗證。令 $\hat{x}_{1:K}$ 為 draft 候選、$v_r=\arg\max_x q_r(x)$ 為續接位置 $r$ 的驗證器 token,我們接受最長相等前綴 $m$。decoder 提交 $\hat{x}_{1:m}$ 後接 $v_{m+1}$。不匹配時,用驗證器的貪婪 token 取代第一個被拒 draft、丟棄其餘;當 $K$ 個候選全匹配則額外提交 $v_{K+1}$ 作為 bonus token。不用隨機接受或殘差分布重取樣。因此解碼無損:對任何接受樣態,提交序列都等於驗證器自己會產出的貪婪序列。
4 訓練資料
訓練資料處理兩件事:標籤品質、以及獎勵函數覆蓋的結構之可得性。我們組裝並清洗公開語料的混合、在開放資料稀疏或難以評分處合成頁面,並為每個樣本配一個任務指令;收集由模型錯誤分析引導。
表 2:Jina-OCR-v1 訓練資料組成。
| 家族 | 代表來源 | 角色 |
|---|---|---|
| Born-digital PDFs | olmOCR-mix, FinePDFs, LightOnOCR, DoclingMatrix | 核心解析 |
| 合成頁面 | olmOCR-synthmix, HTML renderings | 版面多樣性 |
| 公式與表格 | LaTeX-OCR, UniMER, MMTab, SynthChartNet | 結構保真 |
| 歷史與退化 | Europeana, Library of Congress, NARA, NCSE, jaWildText | 穩健性 |
| 多語 | FinePDFs, Chinese PDFs | 語言覆蓋 |
| 商業與表單 | Invoices, CommonForms, RVL-CDIP | 關鍵資訊抽取 |
| VQA 與 grounding | DocLocal4K, VDR | 輔助視覺覆蓋 |
| 獎勵覆蓋合成 | JinaOCRSynth(公式與表格) | GRPO 獎勵覆蓋 |
4.1 開放資料收集與清洗
混合取自多個公開文件 OCR 語料(見表 2)。對困難的真實頁面,加入歷史與退化來源(Europeana 報紙、國會圖書館轉錄、NARA 撫恤檔、NCSE 期刊、日文野生文字樣本)。清洗結合兩個過濾器:規則式——檢測器檢視長標籤最後 1,024 字元中的精確/容截斷/含間隙重複,丟棄找到的退化迴圈;也丟棄完全重複影像與短於 64 token 的樣本。模型式——用 Qwen3.5-122B-A10B 重標歷史報紙/檔案/表單等難源,用 Qwen3-VL-30B-A3B 稽核生成或雜訊樣本。收集由「分類器驅動檢索」的錯誤分析引導。
4.2 有針對性的合成
合成有三個目的。其一,版面多樣性:從 Wikidata 等開放授權來源渲染自足的 HTML 文件(報告、投影片、部落格、資訊圖)。其二,難度:從退化掃描與歷史報紙做偏好對——教師模型精修草稿轉錄、裁判模型對照頁面影像驗證改善,只保留有明確勝者的對。其三,獎勵覆蓋:在自然分布的頁面上,公式與表格專屬獎勵項對多數樣本是空的,故大量 GRPO rollout 沒有結構訊號。JinaOCRSynth 因此把每頁塞滿可評分結構——涵蓋 accent/字型混淆、矩陣環境、macro 慣例、unicode 字形的公式套件,以及源自觀察到失敗(跨欄表頭、密集數值格、轉置或無框版面)的表格套件;每個合成頁附 olmOCR-Bench 式單元測試,讓對應獎勵項適用於它。
4.3 指令覆蓋
訓練 prompt 來自一組任務特定指令:多種格式慣例的整頁解析、表格與公式的元素級轉錄、影像 captioning、文件 VQA、關鍵資訊抽取,並有中英文變體。每個資料集配一個指定 prompt,部分樣本用動態生成的逐圖指令。grounding 與 VQA 作為輔助視覺覆蓋;部署目標是 OCR。
5 後訓練
表 3:Jina-OCR-v1 的四個訓練目標。 對齊 SFT、穩健性 SFT、GRPO 在外迴圈各輪反覆套用(§5.5),FastMTP 在所得驗證器上擬合一次。
| 目標 | 焦點 | 可訓練模組 | 學習率 | 序列長度 |
|---|---|---|---|---|
| 對齊 SFT | 指令、長度外推 | Decoder、projector | $6{\times}10^{-6}$ | 10K, packed |
| 穩健性 SFT | 退化與困難頁 | 視覺塔、decoder | $10^{-6}$ | 10K |
| GRPO | 公式與表格保真 | Decoder、projector | 每 run 而定 | 2–4K completion |
| FastMTP | Draft 頭擬合 | Draft 頭 | $10^{-4}$–$5{\times}10^{-4}$ | 4–8K |
5.1 對齊 SFT 與長度外推
對齊 SFT 讓 base 模型跟隨指令。我們訓練語言模型 decoder、跨模態 projector 與輸出頭(用序列 packing),並凍結視覺塔(SAM、壓縮器、CLIP)。兩個設計選擇:其一,不做長序列訓練就擴展可用位置範圍——base 模型以 8K 情境訓練,而密集報紙需更長輸出,故以 0.4 機率把回應 token 的 position ID 隨機平移一個從 $[1,\min(5L, L_{\mathrm{train}}/4)]$ 抽的偏移(不動 prompt 位置、ID 上限 32,768),讓模型看到超過 packed 長度的 RoPE 位置。其二,平衡混合——跨約 15 個文件領域以階層、領域平衡的比率取樣,讓沒有單一來源主導。
5.2 穩健性微調
穩健性微調把視覺塔以 $10^{-6}$ 解凍、關閉 packing 讓長頁保持完整。結合三個機制:對歷史掃描以 0.25 機率施加光度與幾何腐蝕(小歪斜、正交旋轉、2× 降尺度、椒鹽雜訊、模擬褪色與墨水暈染的形態學侵蝕/膨脹);把標籤改寫向單一格式慣例(如公式個別以 $x+y$,$t+v$ 分隔而非逗號分隔的 $x+y,t+v$),避免格式變異主導損失;並上調歷史報紙掃描、退化文件與公式/表格密集的合成頁的權重,且對每輪外迴圈後錯誤分析找到的失敗模式做過取樣。
5.3 稠密可驗證獎勵下的 GRPO
以結構為焦點的 GRPO run 用 DAPO 變體、每 prompt 八次 rollout。每個獎勵項都由「對照參考轉錄的確定性程式碼」計算,屬 RLVR 設定。兩個 OCR 專屬調適:
- 其一,用 ReMax baseline 取代組正規化 baseline(取樣獎勵減貪婪獎勵,不做獎勵縮放)。OCR rollout 近乎確定,組內獎勵變異常很小,除以它會放大雜訊;貪婪 rollout baseline 避開組統計。
- 其二,關於獎勵本身。二元通過/失敗獎勵稀疏,而分級的部分分數從同樣的確定性檢查提供稠密監督。如表 4,總獎勵是可驗證項的乘積:內容相似度(混合 LaTeX/HTML 的正規化編輯距離)、按資料集選的公式與表格訊號、結構有效性(括號平衡、標籤閉合、表格完整性)、通過的 olmOCR 式單元測試比例、以及反重複與格式項。結構、單元測試與格式項下限 0.2(表格項 0.1)——在乘法組合下,單一失敗檢查會把乘積壓成零、抹掉一個內容正確 completion 的梯度訊號。重複項不設下限,因為退化迴圈是最易膨脹內容分的失敗模式。
表 4:GRPO 的乘法式獎勵組成。(各 run 啟用的項與下限不同)
| 組件 | 訊號 | 角色 |
|---|---|---|
| Content | 混合 LaTeX/HTML 的正規化編輯距離 | 文字保真 |
| Formula | 公式字串匹配 | 公式正確性 |
| Table | TEDS、TEDS-S、表格編輯距離 | 結構還原 |
| 結構有效性 | 括號平衡、標籤閉合、表格完整性 | 良構性 |
| 單元測試 | 通過的 olmOCR 式(存在/順序/數學/表格)測試比例 | 稠密回饋 |
| 重複與格式 | 重複懲罰、HTML 合規 | 退化控制 |
5.4 輔助正則器
部分 SFT 與 GRPO run 對繼承的 MoE decoder 加輔助正則器:序列級負載平衡損失 $L_{\mathrm{bal}}=N\sum_i f_i P_i$、正規化 gate 熵項 $L_{\mathrm{ent}}$、PaLM 的輸出 logit 懲罰 $L_z=\lambda_z\,\mathrm{logsumexp}(z)^2$。在 token 表示開始崩塌處,加一個 margin 為 $m$ 的 SimCTG 式對比正則器 $L_{\mathrm{ctr}}$(鼓勵鄰近 token 狀態分離,$s_{ij}=\cos(h_i,h_j)$),沿用 ReaderLM-v2 的正則化。
5.5 代理式檢查點合併與資料整理
對齊、穩健性與 GRPO 產出一池候選檢查點(資料混合、獎勵組成、正則化各異),在外迴圈各輪反覆套用:訓練候選 → 合併 → 診斷失敗 → 整理資料。draft 頭在迴圈選定驗證器後才擬合一次。
合併是一個搜尋問題:在固定評測預算下,一個代理平行評估合併配置、用單元測試與編輯距離檢查打分、由先前結果規劃下一批合併。所選合併接著做錯誤分析,高價值失敗模式由 §4 的方法檢索或合成,餵入下一輪對齊、穩健性微調或 GRPO。
圖 4(外訓練迴圈,文字描述):SFT 與 GRPO run 產出候選檢查點 → 代理在評測預算下合併 → 錯誤分析驅動下一個資料混合。FastMTP 在迴圈後對選定的驗證器擬合。
5.6 Draft 頭目標
外迴圈選定驗證器後,此階段對該凍結模型擬合 draft 頭(用 §3.1 的 state-shift 對齊)。對每個有效回應列 $p$,定義 $P_p^{(k)}=\mathrm{softmax}(\ell_p^{(k)})$ 與 detach 的驗證器分布 $Q_p=\mathrm{softmax}(W_{\mathrm{lm}}b_{p+1})$。每個遞迴深度用同一列對齊的 token 與狀態目標。以 $\alpha_k=\beta^{k-1}/\sum_r\beta^{r-1}$、$\beta=0.6$,目標為交叉熵項與分布對齊項的加權和($L=0.85\sum_k\alpha_k L_{\mathrm{CE}}^{(k)}+0.45\sum_k\ldots$,驗證器狀態與分布 detach)。可訓練參數為 $B_\theta, P, N_e, N_h$;驗證器骨幹固定,draft 嵌入、輸出 norm 與 LM head 綁定到驗證器權重、不更新。
6 評測
6.1 基準與指標
評測解析品質與服務效能。olmOCR-Bench 以單元測試式檢查(文字存在、閱讀順序、數學、表格)給頁面級解析評分,回報含 headers/footers 類別的官方 overall。OmniDocBench v1.6 回報文字編輯距離、公式 CDM、表格 TEDS/TEDS-S、閱讀順序編輯距離,overall 平均文字分、公式 CDM 與表格 TEDS。baseline 為近期通用 VLM 與有公開數字的特化 OCR 模型。
6.2 解析品質
表 5:olmOCR-Bench 比較。 Params 給總參數量,MoE 給激活參數量;每欄最佳粗體。
| 模型 | Params | ArXiv | OldScans-Math | Tables | OldScans | Multi-col | LongTiny | Hdr/Ftr | Base | Overall |
|---|---|---|---|---|---|---|---|---|---|---|
| 通用 VLM | ||||||||||
| Gemini 3 Flash | – | 80.1 | 73.6 | 64.6 | 45.8 | 75.3 | 90.3 | 27.4 | – | – |
| Qwen3-VL-235B | 235B/22B | 88.4 | 81.2 | 86.7 | 49.6 | 85.9 | 88.9 | 33.6 | – | – |
| 特化 OCR 模型 | ||||||||||
| DeepSeek-OCR | 3B/570M | 77.5 | 74.5 | 77.3 | 33.1 | 67.3 | 83.0 | 96.1 | 99.3 | 76.0 |
| dots.mocr | 3B | 85.9 | 85.5 | 90.7 | 48.2 | 85.3 | 81.6 | 94.0 | 99.7 | 83.9 |
| olmOCR-2 | 8B | 82.9 | 82.1 | 84.3 | 48.3 | 84.3 | 81.4 | – | 99.7 | 82.4 |
| LightOnOCR-2 | 1B | 89.6 | 85.6 | 89.0 | 42.2 | 84.8 | 91.4 | 19.7 | 99.6 | 83.2 |
| chandra-ocr-2 | 4B | 86.9 | 89.1 | 92.1 | 51.1 | 82.1 | 93.7 | 91.4 | 99.9 | 85.8 |
| Jina-OCR-v1 | 3B/570M | 86.1 | 82.3 | 88.8 | 42.6 | 85.5 | 93.2 | 88.7 | 99.9 | 83.4 |
Jina-OCR-v1 達 83.4 overall,在有公開 overall 的特化模型中排第三(次於 chandra-ocr-2 的 85.8 與兩階段的 dots.mocr 83.9),領先 LightOnOCR-2 (83.2) 與 8B 的 olmOCR-2 (82.4)。它比它後訓練的骨幹 DeepSeek-OCR (76.0) 高 7.4 分——這是表中唯一隔離出本文後訓練配方效果的比較。
表 6:OmniDocBench v1.6 比較。(RO=閱讀順序;每欄最佳粗體。編輯距離越低越好↓)
| 方法 | Params | Overall↑ | Text Edit↓ | Formula CDM↑ | Table TEDS↑ | TEDS-S↑ | RO Edit↓ |
|---|---|---|---|---|---|---|---|
| 通用 VLM | |||||||
| Gemini 3 Flash | – | 92.62 | 0.066 | 95.16 | 89.29 | 93.51 | 0.172 |
| Qwen3-VL-235B | 235B/22B | 89.78 | 0.063 | 92.55 | 83.07 | 86.75 | 0.166 |
| 特化 OCR 模型 | |||||||
| DeepSeek-OCR-2 | 3B/570M | 90.25 | 0.050 | 91.84 | 83.89 | 87.75 | 0.144 |
| HunyuanOCR-1.5 | 1B | 94.74 | 0.039 | 94.50 | 93.67 | 94.71 | 0.129 |
| PaddleOCR-VL-1.6 | 0.9B | 96.34 | 0.033 | 97.53 | 94.76 | 97.10 | 0.128 |
| Jina-OCR-v1 | 3B/570M | 91.14 | 0.046 | 93.28 | 84.68 | 89.01 | 0.142 |
在 OmniDocBench 上 Jina-OCR-v1 達 91.14 overall,在有公開 v1.6 分的特化模型中排第三(次於 PaddleOCR-VL-1.6 的 96.34 與 HunyuanOCR-1.5 的 94.74),領先 DeepSeek-OCR-2 (90.25) 與大得多的 Qwen3-VL-235B (89.78);相對 DeepSeek-OCR-2 的優勢在每一欄都成立(含公式 CDM 93.28 對 91.84)。
6.3 服務吞吐
表 7:OCR 模型在 olmOCR-Bench 上的服務效率(1,403 頁,單張 A100 SXM4 40GB,並發 32)。Score 為 olmOCR-Bench overall。
| 模型 | Score↑ | Pages/s↑ | Output tok/page↓ | Output tok/s↑ |
|---|---|---|---|---|
| chandra-ocr-2 | 85.8 | 0.38 | 1917 | 730 |
| dots.mocr | 83.9 | 0.55 | 1711 | 934 |
| Surya OCR 2 | 83.3 | 1.05 | 3568 | 3760 |
| LightOnOCR-2 | 83.2 | 1.33 | 1208 | 1606 |
| Infinity-Parser-7B | 82.5 | 0.64 | 1066 | 680 |
| olmOCR-2 | 82.4 | 1.22 | 1128 | 1374 |
| MinerU2.5-Pro-1.2B | 82.2 | 1.28 | 1548 | 1982 |
| DeepSeek-OCR | 76.0 | 2.10 | 1366 | 2871 |
| GLM-OCR | 75.2 | 1.21 | 1054 | 1272 |
| MinerU2.5-1.2B | 75.2 | 1.43 | 1563 | 2231 |
| PaddleOCR-VL-1.6 | – | 0.33 | 1048 | 344 |
| HunyuanOCR-1.5 | – | 0.83 | 1058 | 873 |
| DeepSeek-OCR-2 | – | 1.50 | 1389 | 2083 |
| Jina-OCR-v1 | 83.4 | 2.57 | 1085 | 2792 |
token 吞吐不單獨決定頁面吞吐:Surya OCR 2 的 tok/s 最高 (3,760) 卻每頁輸出 3,568 token,而 Jina-OCR-v1 每頁只輸出 1,085 token、更快完成頁面。兩者關係為 pages/s = (tok/s)/(tok/page),且輸出長度與解析品質無關($r{=}{+}0.22$)——因此精簡度可獨立最佳化,Jina-OCR-v1 以「有競爭力的 token 吞吐 + 所有得分 83 以上系統中最短的輸出」達到最高頁面吞吐。
6.4 推測式解碼
表 8:Jina-OCR-v1 帶 FastMTP 在低預算 GPU(NVIDIA L4,vLLM 0.20.1,batch=1)上的解碼效率。 Speedup $S$ 相對同執行模式的 $k{=}0$ 基線;$\tau$ 為每推測步提交的平均 token 數(含 bonus)、$c=\tau/S$ 為一推測步的成本(以一自迴歸步為單位)。
| 模式 | k | Output tok/s↑ | Speedup S↑ | 接受率 | τ | c↓ |
|---|---|---|---|---|---|---|
| Eager | 0 | 42.7 | 1.00× | – | – | 1.00 |
| 1 | 64.0 | 1.50× | 82.6% | 1.83 | 1.22 | |
| 2 | 77.9 | 1.82× | 69.1% | 2.38 | 1.30 | |
| 3 | 83.1 | 1.95× | 57.6% | 2.73 | 1.40 | |
| Graph | 0 | 158.3 | 1.00× | – | – | 1.00 |
| 1 | 185.6 | 1.17× | 82.9% | 1.83 | 1.56 | |
| 2 | 183.8 | 1.16× | 69.3% | 2.38 | 2.05 | |
| 3 | 172.9 | 1.09× | 57.9% | 2.74 | 2.51 |
FastMTP 在 $k{=}3$ 達 eager 自迴歸基線的 1.95×。CUDA graph 把該基線從 42.7 提到 158.3 tok/s(3.71×),對它最佳深度是 $k{=}1$ 的 1.17×。接受行為與模式無關($k{=}3$ 時 $\tau$ eager 2.73、graph 2.74)。條件接受率從第一 draft 位置的 0.83 降到第三的 0.65,故每增一深度貢獻遞減而成本固定——eager 最佳深度 $k{=}3$、CUDA graph 下 $k{=}1$。
7 結論
我們提出 Jina-OCR-v1,一個建於 DeepSeek-OCR 壓縮視覺、3B-MoE 骨幹之上的端到端文件解析器,配一個遞迴共享的 FastMTP draft 頭與稠密可驗證獎勵下的後訓練。它在 olmOCR-Bench 得 83.4、OmniDocBench v1.6 得 91.14,並在我們量測的系統中是每秒頁數最快者。頁面吞吐與 token 吞吐對系統的排序不同——因為解碼較快的系統也輸出較長的轉錄,而輸出長度與解析品質無關,故精簡度可獨立最佳化。在 NVIDIA L4 這類低預算 GPU 上,FastMTP 頭把解碼速度近乎翻倍,且接受行為對驗證器的執行模式不變。
參考文獻
【依原文順序(作者-年份制,保留原文英文)。】
- Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, and William Brandon. Hydra: Sequentially-dependent draft heads for medusa decoding, 2024.
- Léo Appourchaux, Noé Brandolini, Paul Lemaistre, and André-Louis Rochet. Visual document retrieval (VDR) multi-domain dataset. https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval , 2025.
- Shuai Bai, Keqin Chen, Xuejing Liu, et al. Qwen2.5-VL technical report. arXiv preprint arXiv:2502.13923 , 2025.
- Yong Yi Bay and Kathleen A. Yearick. GRPO, Dr. GRPO, and DAPO are three operations on one number: The group-standard-deviation identity, 2026.
- Lukas Blecher, Guillem Cucurull, Thomas Scialom, and Robert Stojnic. Nougat: Neural optical understanding for academic documents. arXiv preprint arXiv:2308.13418 , 2023.
- Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. In International Conference on Machine Learning (ICML) , 2024.
- Yuxuan Cai, Xiaozhuan Liang, Xinghua Wang, Jin Ma, Haijin Liang, Jinwen Luo, Xinyu Zuo, Lisheng Duan, Yuyang Yin, and Xi Chen. FastMTP: Accelerating LLM inference with enhanced multi-token prediction. arXiv preprint arXiv:2509.18362 , 2025.
- Jian Chen, Yesheng Liang, and Zhijian Liu. DFlash: Block diffusion for flash speculative decoding. arXiv preprint arXiv:2602.06036 , 2026.
- Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research , 24(240):1–113, 2023.
- Cheng Cui, Ting Sun, Man Lin, et al. PaddleOCR 3.0 technical report. arXiv preprint arXiv:2507.05595 , 2025.
- Damai Dai, Chengqi Deng, Chenggang Zhao, et al. DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066 , 2024.
- Shuaiqi Duan, Yadong Xue, Weihan Wang, et al. GLM-OCR technical report. arXiv preprint arXiv:2603.10910 , 2026.
- Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. Break the sequential dependency of LLM inference using lookahead decoding. In International Conference on Machine Learning (ICML) , 2024.
- GLM 4.5 Team. GLM-4.5: Agentic, reasoning, and coding (ARC) foundation models, 2025.
- Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, and Gabriel Synnaeve. Better & faster large language models via multi-token prediction. In International Conference on Machine Learning (ICML) , 2024.
- Adam W. Harley, Alex Ufkes, and Konstantinos G. Derpanis. Evaluation of deep convolutional nets for document image classification and retrieval. In International Conference on Document Analysis and Recognition , 2015.
- Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, et al. mPLUG-DocOwl 1.5: Unified structure learning for OCR-free document understanding. arXiv preprint arXiv:2403.12895 , 2024.
- Hugging Face. Finepdfs. https://huggingface.co/datasets/HuggingFaceFW/finepdfs , 2025.
- INF Team. Infinity-Parser2 technical report. arXiv preprint arXiv:2607.07836 , 2026.
- Alexander Kirillov, Eric Mintun, Nikhila Ravi, et al. Segment anything. arXiv preprint arXiv:2304.02643 , 2023.
- Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luca Soldaini, Noah A. Smith, Yizhong Wang, Pradeep Dasigi, and Hannaneh Hajishirzi. Tülu 3: Pushing frontiers in open language model post-training, 2025.
- Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning , 2023.
- Gengluo Li, Xingyu Wan, Shangpin Peng, et al. HunyuanOCR-1.5: Making lightweight OCR VLMs faster and better. arXiv preprint arXiv:2607.04884 , 2026.
- Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE: Speculative sampling requires rethinking feature uncertainty. arXiv preprint arXiv:2401.15077 , 2024.
- Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. EAGLE-3: Scaling up inference acceleration of large language models via training-time test. arXiv preprint arXiv:2503.01840 , 2025a.
- Zhang Li, Yuliang Liu, Qiang Liu, et al. MonkeyOCR: Document parsing with a structure-recognition-relation triplet paradigm. arXiv preprint arXiv:2506.05218 , 2025b.
- Ziniu Li, Tian Xu, Yushun Zhang, et al. ReMax: A simple, effective, and efficient reinforcement learning method for aligning large language models. arXiv preprint arXiv:2310.10505 , 2023.
- Aixin Liu, Bei Feng, Bin Wang, et al. DeepSeek-V2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434 , 2024a.
- Aixin Liu, Bei Feng, Bing Xue, et al. DeepSeek-V3 technical report. arXiv preprint arXiv:2412.19437 , 2024b.
- Yuliang Liu, Zhang Li, Ziyang Zhang, et al. MonkeyOCRv2: A visual-text foundation model for document AI. arXiv preprint arXiv:2607.11562 , 2026.
- Living with Machines Consortium. Living with Machines: Historical newspapers (ln_newspapers). https://huggingface.co/datasets/biglam/ln_newspapers , 2023.
- llm-jp. jaWildText: Japanese wild-text OCR samples. https://huggingface.co/datasets/llm-jp/jawildtext , 2024.
- Ahmed Nassar, Andrey Marafioti, Matteo Omenetti, et al. SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion. arXiv preprint arXiv:2503.11576 , 2025.
- Linke Ouyang, Yuan Qu, Hongbin Zhou, et al. OmniDocBench: Benchmarking diverse PDF document parsing with comprehensive annotations. In Proceedings of the Computer Vision and Pattern Recognition Conference , pages 24838–24848, 2025.
- Jake Poznanski, Akshita Rangapur, Jon Borchardt, et al. olmOCR: Unlocking trillions of tokens in PDFs with vision language models. arXiv preprint arXiv:2502.18443 , 2025.
- Jake Poznanski, Kyle Lo, and Luca Soldaini. The olmOCR project: Building fully open OCR using VLMs. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) , pages 626–635, 2026.
- Qwen Team. Qwen3.5. https://huggingface.co/Qwen/Qwen3.5-122B-A10B , 2026.
- Alec Radford, Jong Wook Kim, Chris Hallacy, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning , pages 8748–8763, 2021.
- Rednote HiLab. dots.ocr. https://github.com/rednote-hilab/dots.ocr , 2025.
- RevolutionCrossroads. Chronicling america: Historic american newspapers, 1770–1810. https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810 , 2024a.
- RevolutionCrossroads. Nara revolutionary war pension files. https://huggingface.co/datasets/RevolutionCrossroads/nara_revolutionary_war_pension_files , 2024b.
- Alexander Samarin, Sergei Krutikov, Anton Shevtsov, Sergei Skvortsov, Filipp Fisin, and Alexander Golubev. LK losses: Direct acceptance rate optimization for speculative decoding. arXiv preprint arXiv:2602.23881 , 2026.
- Zhihong Shao, Peiyi Wang, Qihao Zhu, et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300 , 2024.
- Idan Shenfeld, Jyothish Pari, and Pulkit Agrawal. RL’s razor: Why online reinforcement learning forgets less. In International Conference on Learning Representations (ICLR) , 2026.
- Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. Blockwise parallel decoding for deep autoregressive models. In Advances in Neural Information Processing Systems (NeurIPS) , 2018.
- Jianlin Su, Yu Lu, Shengfeng Pan, et al. RoFormer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864 , 2021.
- Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. A contrastive framework for neural text generation. In Advances in Neural Information Processing Systems , 2022.
- Ting Sun, Cheng Cui, Yuning Du, and Yi Liu. Pp-doclayout: A unified document layout detection model to accelerate large-scale data construction, 2025.
- Said Taghadouini, Adrien Cavaillès, and Baptiste Aubertin. LightOnOCR: A 1B end-to-end multilingual vision-language model for state-of-the-art OCR. arXiv preprint arXiv:2601.14251 , 2026.
- UCL Centre for Digital Humanities. Nineteenth-Century Serials Edition (NCSE) v2.0: OCR-processed 19th-century English newspapers. https://rdr.ucl.ac.uk/articles/dataset/28381610 , 2024.
- Bin Wang, Fan Wu, Linke Ouyang, Zhuangcheng Gu, Rui Zhang, Renqiu Xia, Botian Shi, Bo Zhang, and Conghui He. Image over text: Transforming formula recognition evaluation with character detection matching. arXiv preprint arXiv:2409.03643 , 2024a.
- Bin Wang, Chao Xu, Xiaomeng Zhao, et al. MinerU: An open-source solution for precise document content extraction. arXiv preprint arXiv:2409.18839 , 2024b.
- Feng Wang, Zesheng Shi, Bo Wang, Nan Wang, and Han Xiao. ReaderLM-v2: Small language model for HTML to Markdown and JSON. arXiv preprint arXiv:2503.01151 , 2025.
- Longwen Wang, Yirui Liu, Xiaohui Hu, and Yuankai Fan. Beyond binary: Turning partial success into dense verifiable rewards for reinforcement learning in code generation, 2026.
- Haoran Wei, Chenglong Liu, Jinyue Chen, et al. General OCR theory: Towards OCR-2.0 via a unified end-to-end model. arXiv preprint arXiv:2409.01704 , 2024.
- Haoran Wei, Yaofeng Sun, and Yukun Li. DeepSeek-OCR: Contexts optical compression. arXiv preprint arXiv:2510.18234 , 2025.
- Qiying Yu, Zheng Zhang, Ruofei Zhu, et al. DAPO: An open-source LLM reinforcement learning system at scale. arXiv preprint arXiv:2503.14476 , 2025.
- Mingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She, Zheng Lin, Wenbin Jiang, and Weiping Wang. Multimodal table understanding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics , pages 9102–9124, 2024.
- Xu Zhong, Elaheh Tangborn, and Jayant Saunshi. Image-based table recognition: Data, model, and evaluation. In European Conference on Computer Vision , 2020.
術語對照表
| English | 繁體中文 |
|---|---|
| Document parsing | 文件解析 |
| Compressed-Vision Encoder | 壓縮視覺編碼器 |
| Mixture-of-Experts (MoE) | 專家混合 |
| Speculative Decoding | 推測式解碼 |
| Multi-Token Prediction (MTP) | 多 token 預測 |
| FastMTP / draft block | FastMTP / draft block(草擬區塊) |
| Draft head | draft 頭(草擬頭) |
| Verifier | 驗證器 |
| Greedy Verification | 貪婪驗證 |
| Lossless | 無損 |
| Acceptance rate | 接受率 |
| Bonus token | bonus token(額外 token) |
| Dense Verifiable Rewards | 稠密可驗證獎勵 |
| RLVR (RL with Verifiable Rewards) | 可驗證獎勵強化學習 |
| GRPO / DAPO | 組相對策略最佳化 / DAPO |
| ReMax baseline | ReMax 基線 |
| Partial credit | 部分分數 |
| Degeneration loop | 退化迴圈 |
| Dynamic resolution / Gundam mode | 動態解析度 / Gundam 模式 |
| Alignment SFT | 對齊 SFT |
| Robustness fine-tuning | 穩健性微調 |
| Length extrapolation | 長度外推 |
| Load-balancing loss | 負載平衡損失 |
| SimCTG contrastive regularizer | SimCTG 對比正則器 |
| TEDS / CDM | TEDS / CDM(表格與公式評測指標) |
| Page throughput | 頁面吞吐 |
| Pixels per visual token | 每視覺 token 像素數 |