發明一個資料集:以零種子測量資料集生成能力
原始論文:Invent a Dataset: Measuring dataset generation abilities with zero seed 作者:Singh, Shivalika、Djurisic, Andrija、Onilude, Gbemileke、Roy, Sudip、Hooker, Sara arXiv ID:2610.01674v1 日期:1 Oct 2026 標籤:LLM 開源模型 合成資料 數據生成 數據增強 數據集生成 零樣本
【譯註:本篇譯文由自動化管線產出:先取 PDF 版面結構,再以 LLM 逐段翻譯並依全站術語詞表統一譯名(模型本身不對外公開資訊,僅說明為 LLM);未經逐句人工校對。】
目錄
- 1 引言
- 2 實驗設置
- 3 結果
- 4 相關研究
- 5 結論
- 限制 ---
- 附錄
- 3 品質與相關性裁判模板
- 3. System Message
- 提示品質評分準則
- 規模
- 特別案例
- 完成品質評分準則
- 規模
- 特別案例
- 完整裁判提示
- 目標規格
- 軸 1 — relevance_score
- 3 relevance_score scale
- ### 決策
- 3. 僅回傳此 JSON 物件,且不得包含其他內容
- 輸出格式
Shivalika Singh¹, Andrija Djurisic¹, Gbemileke Onilude¹, Sudip Roy¹, and Sara Hooker¹
¹Adaptation
建立資料集一直是 AI 開發中最耗費人力且最脆弱的環節。在本技術報告中,我們聚焦於實務工作者面臨的最極端但同時也最普遍的情境:零資料機制。在此情境下,實務工作者沒有任何可用於學習目標能力的資料。我們推出 INVENT-A-DATASET,這是一個基於提示(Prompt)的系統,能從資料集描述直接生成真實且具大規模的後訓練(Post-training)資料集。我們將 INVENT-A-DATASET 與五個前沿模型 API 進行評估,包括 Anthropic ✶、Google ♦、OpenAI 🧨、DeepSeek 👥 以及 Zai 🧗。在八種任務類型與最多 20K 樣本的資料集規模下,INVENT-A-DATASET 在品質(17% 相對增益)與多樣性(19% 相對增益)兩方面均顯著優於其他工具。其多樣性優勢會隨著訓練資料集規模的擴大而更加顯著(從 200 樣本時的持平狀態,提升至 20K 樣本時的 37% 相對增益)。這轉化為顯著的下游訓練增益,進而產生效能更優異的後訓練模型。相較於其他生成器微調(fine-tune),Invent API 微調在不同後訓練模型架構下,始終保持領先地位。
1 引言
資料集是推動突破的原料。人工智慧領域的進展日益仰賴高品質的訓練資料來持續挑戰模型 (Kulikov et al., 2026)。缺乏資料存取權,也是決定誰能主導人工智慧發展的重要因素之一 (Longpre et al., 2024)。然而,要取得高品質且數量充足的資料,以針對新模型能力進行訓練,仍面臨極大的挑戰。人工標記的訓練資料蒐集成本高昂,且過去一直被視為世界靜態的代表性 (Kulikov et al., 2026; Salazar et al., 2026; Singh et al., 2025; 2024a; Boubdir et al., 2023b)。蒐集額外資料需要耗費大量時間,導致資料集與我們身處的豐富且不斷演進的環境相去甚遠 (Roh et al., 2019)。
在本文技術報告中,我們聚焦於此資料鴻溝中最極端的一端:即「零訓練資料機制」(zero training data regime),亦即缺乏目標能力所需資料的狀態。我們嚴肅看待此情境,因為這也是現實世界從業者最常面臨的挑戰。雖然已有數種解決方案針對此零資料機制提出,但皆未能同時保證最終資料集的「多樣性」與「品質」。資料集檢索代理人雖利用網路搜尋來尋找相關資料集,但此方法仍存在限制。
- 通訊作者。

圖 1:多樣性與品質的權衡。以 DCScore($τ = 0.05$)衡量四個外部模型與 Invent API 的品質(x 軸)與多樣性(y 軸),兩項分數皆為在無約束設定下,針對八個 5K 樣本資料集取平均值。兩軸皆為越高越好,右上角為最佳。Invent API 同時達成了高品質與高多樣性。
僅僅回傳資料集,並不保證其格式符合 AI 訓練需求 (Lin et al., 2025; Chapman et al., 2020b; Li et al., 2025a)。此外,取得的資料集也不太可能滿足從業人員的所有獨特需求,例如格式、語言覆蓋率或輸出語調。這限制了導入新屬性或針對任務特定指標進行明確最佳化的可行性。
另外,實務工作者通常會使用現有的模型來產生與其能力描述相關的合成資料。然而,大多數合成資料技術都假設你已經具備一個種子池(seed pool)提示詞,或是已經擁有需要改進的資料集(Wang et al., 2023a; Xu et al., 2024)。若缺乏這些種子提示詞,一項重大挑戰在於如何在成功進行後訓練(Post-training)的同時,將資料集擴展至所需規模,並維持語料庫的多樣性(Dohmatob et al., 2024; Briesch et al., 2023; Shumailov et al., 2023; Bertrand et al., 2024; Guo et al., 2024)。大型語言模型(LLMs)經常面臨輸出多樣性低的問題(Holtzman et al., 2019; Xu et al., 2022)。在缺乏多樣化提示詞的情況下,採用標準最大化解碼(maximization-based decoding)的模型已知會陷入冗餘的連續重複。此外,為了訓練目的而蒸餾合成資料時,通常會受到授權限制:許多專有(Proprietary)的頂尖模型在其服務條款(terms of service)中明確限制在其實作輸出上進行訓練(Wiggers, 2025)。
在本研究中,我們提出 INVENT-A-DATASET 以解決上述多項痛點。INVENT-A-DATASET 實現了自主且需求驅動的資料策展,能將單一資料集描述轉化為具備真實性且已準備好用於 AI 訓練的資料 (Adaption Labs, 2026)。我們以寬鬆授權釋出 INVENT-A-DATASET,該授權明確允許在訓練中進行下游商業用途。我們針對 Anthropic、Google、OpenAI、DeepSeek 與 Zai 等五個前沿模型 API,在樣本數最多達 20K 的多樣化資料集規模下,評估 INVENT-A-DATASET 的品質、多樣性與資料集相關性。我們的研究發現如下:
- Invent API 在品質方面顯著優於其他模型(相較於最強基線,品質提升了 17% 的相對增益),同時也能產生最多樣化的樣本(在 2k 樣本下,多樣性提升了 19%)。 2. 這種在多樣性上的穩定優勢會隨著規模擴大而更加顯著。與大多數基線模型不同,後者的多樣性在資料集請求從 2K 增加到 20K 樣本時會停滯或下降(Claude-Opus-5 下降 15%、GLM-5.3 下降 8%、GPT-5.6-sol 下降 7%、DeepSeek-V4-pro 提升 0.4%),而 Invent API 則隨著規模擴大變得更加多樣(提升 7%),同時維持穩定的品質。 3. 在資料集請求中加入語氣與輸出格式等約束,會使所有評估模型的輸出多樣性降低 15% 至 22%。相比之下,Invent API 在受約束與更具挑戰性的不受約束查詢中皆能維持其多樣性,在 20K 樣本的受約束查詢中,其表現甚至比第二名的 GLM-5.3 高出驚人的 58%。 4. 這些品質提升會轉化為下游後訓練(Post-training)的品質優勢。我們獨立進行後訓練,並針對不同資料變體對微調模型的生成結果進行排名。在 Llama-3.3-70B 的測試提示中,Invent API 微調模型在 54% 的情況下排名第一;在 Gemma-4-31B-it 的測試提示中,則在 41% 的情況下排名第一。其表現優於未經微調的基礎模型,以及在 Claude-Opus-5 與 GLM-5.3 生成資料上微調的模型。我們的結果顯示,現在已可行透過零資料機制(zero data regime)自動鎖定新的能力。過去,由於收集與策展資料的成本高昂,難以「即時」(on-the-fly)調整訓練集以增加覆蓋率或提升任務多樣性。
INVENT-A-DATASET 是邁向持續優化訓練集的重要一步,這些訓練集能隨任務與目標能力演進而進化。
2 實驗設置
2.1 評估設置
Invent 是一個專門用於從提示產生多元化資料集的 API 端點。我們將 Invent API 與多種專有及開放權重模型進行基準測試。針對每個模型、目標資料集大小與資料集查詢,我們衡量其產生高品質且多元化資料集的能力。我們評估了八個具代表性的資料集查詢,旨在涵蓋 1) 任務類型、2) 語言以及 3) 從客戶支援到醫療問答等多元領域。此外,針對其中五個資料集,我們在 tone constraints.json 格式與多選輸出限制等額外格式約束下,設計了更具挑戰性的版本。這些設計旨在反映實務界對於訓練資料集的常見需求(Zhou et al., 2023; Jiang et al., 2024; Wen et al., 2024)。我們將這兩組評估查詢分別稱為不受約束查詢與受約束查詢。三個資料集的評估查詢包含於表 8,其餘則詳見附錄 A.3。
所考慮的模型。我們的基準測試集包含 Anthropic 的 Claude Opus 5 (Anthropic, 2026)、OpenAI 的 GPT-5.6 Sol (OpenAI, 2026) 以及 Google 的 Gemini-3.1-pro 等專有模型。
| Dataset | Unconstrained Query | Constrained Query |
|---|---|---|
| African QA | Dataset containing question-and-answer pairs covering diverse topics related to African research, literature, history, and current events. | Dataset containing multi-turn conversations between a teacher and a student on an exam, covering diverse topics related to African research, literature, history, and current events. Prompts must use OpenAI format while responses should be plain text. |
| Legal QA | Dataset containing question-answer pairs focusing on legal advice, factual explanations, and comparative analysis across diverse topics such as copyright law, government types, and liability. | Dataset containing question-answer pairs focusing on legal advice, factual explanations, and comparative analysis across diverse topics such as copyright law, government types, and liability. Questions should offer four choices; answers are limited to A, B, C, or D with no additional text. |
| Serbian Car Ads | A dataset containing a collection of used car advertisements in Serbian, featuring detailed descriptions of vehicle conditions, service histories, and equipment packages. | A dataset containing a collection of summaries of used car advertisements in Serbian, featuring detailed descriptions of vehicle conditions, service histories, and equipment packages. Each prompt is an ad summary, and each response is a JSON object containing vehicle_description, service_history, and equipment. |
表 3:用於產生資料集的單一請求查詢。這些查詢旨在涵蓋 1) 任務類型、2) 語言以及 3) 領域的多樣性。我們針對不受約束與受約束版本的查詢進行評估,其中受約束的查詢需滿足額外要求,例如 1) 語調約束、2) JSON 格式、3) 多選輸出限制。受約束查詢與原始版本之間的差異已標示出來。
(Google DeepMind, 2026). 我們選擇的專有 API 服務條款禁止將其資料輸出用於訓練。然而,我們仍納入這些模型,旨在評估效能上限,以作為校準的參考基準。此外,許多實務工作者似乎仍依賴這些模型進行合成資料生成,即便此舉已違反規定。我們也納入了開放權重模型,例如 Z.ai 的 GLM-5.3 (Z.ai, 2026) 與 DeepSeek AI 的 DeepSeek-V4-Pro-0813 (DeepSeek-AI, 2026),其授權條款允許將輸出用於下游訓練。在本次實驗中,我們透過 OpenRouter API (OpenRouter, 2026) 存取這些模型。
Generation configurations. We kept all model hyperparameters at their default values (temperature 1.0 where possible to specify, reasoning disabled). INVENT-A-DATASET returns a dataset in a single call, however in practice we observed that other models would rarely return the required number of samples if only a single call was used. At scale, single call requests often led to significant diversity collapse. While we continued to benchmark INVENT-A-DATASET using only a single call, we evaluated a set of different approaches on the other models benchmark : (i) one-by-one, prompting the model to produce a single sample per request (in contrast to requesting entire dataset in one query) (ii) batch of 100, producing samples in fixed batches of N=100 (iii) max batch size, producing as many samples as possible in a single request depending on model's output token limit which was set to 128K across all models. As shown in Figure 2, diversity collapses almost entirely (≈0.00–0.03 across all four models) for the one-by-one approach, while max batch strategy is consistently the most diverse. Based on these results, we selected max batch as the default generation approach.

圖 2:不同生成策略下的多樣性。我們比較了從單一查詢請求(醫療問答)生成資料集的三種模型:Gemini 3.1 Pro、GPT-5.6 Sol 與 GLM-5.3。這些策略的差異在於每次請求產生的樣本數量:一次產生一個樣本(one-by-one)、固定批次大小為 100,或是盡可能在模型的最大輸出 Token 限制內產生樣本(batch_size max)。y 軸顯示 DCScore($τ = 0.1$),數值越高代表多樣性越佳。多樣性通常會隨著批次大小增加而提升,而逐一生成樣本會導致所有模型出現近乎完全的多樣性塌陷($≤ 0.03$)。
我們採用此方法進行後續所有實驗,除非另有說明。
資料集規模化。在實務上,取得充足資料的核心目標在於訓練模型。因此,我們關注於產生足夠規模且具可行性的資料,以供訓練使用。針對每個查詢與模型生成器,我們在四個不同的規模下產生資料:200、2000、5000 與 20,000 個樣本。較大的規模能讓我們測量品質與多樣性如何隨著實務人員對後訓練目的提出更大資料集的需求而產生變化。
下游後訓練設定。為了測試使用 Invent API 是否能提升下游模型效能,我們針對三種模型(Llama-3.3-70B-Instruct、Gemma-4-31B-it 與 Qwen3.5-9B)獨立重複相同的程序,這些模型分別屬於不同的模型家族與規模。我們使用 LoRA SFT 在每個生成器產出的 20K 醫療問答資料變體上對其進行後訓練。我們不手動調整訓練配方,而是使用 Autoscientist (Adaption Research Staff, 2026) 在其中一個醫療問答資料變體上執行超參數搜尋。為了控制僅由資料集產生的影響,Autoscientist 提出的配置隨後會套用於每個生成器的資料集。因此,比較中的每個模型都從相同的基準開始,並使用相同的優化器設定、相同的 epoch 數與相同的樣本數進行訓練,因此訓練資料是它們之間唯一的變數。
“高血壓是如何診斷的?哪些血壓數值是關鍵?” “高血壓的診斷應基於正確測量且重複的血壓讀數,而非單次門診數值。患者應安靜坐著五分鐘,背部有支撐、雙腳平放 [...]”
Gemini 3.1 Pro “Dabigatran 屬於哪一類的抗凝血劑?” “Dabigatran 屬於直接凝血酶抑制劑(DTI)。它可逆地結合凝血酶(Factor IIa)的活性位點,防止纖維蛋白原轉化為纖維蛋白。它是首款獲准用於 [...] 的 DOAC。”
DeepSeek V4 Pro “Digoxin 的作用機制為何?” “Digoxin 抑制心肌細胞中的 ATP 酶幫浦,提升細胞內鈉離子濃度,並透過鈉鈣交換器提升細胞內鈣離子濃度 [...]”
GPT 5.6 Sol “心臟病發作有哪些警訊?” “警訊包括胸部出現壓迫感、緊縮感、沉重感或疼痛,有時會擴散至手臂、肩膀、下顎、背部或上腹部。也可能出現呼吸急促、出汗、噁心、頭暈或異常疲勞,且症狀在女性、年長者及糖尿病患者中可能表現非典型 [...]”GLM 5.3 “What are the diagnostic criteria for hypertension?” “Hypertension is diagnosed based on repeated elevated office blood pressure measurements. Under the 2017 ACC/AHA guidelines, stage 1 is systolic 130–139 or diastolic 80–89 mmHg, and stage 2 is 140/90 mmHg or higher. Out-of-office ambulatory or home monitoring is recommended to confirm the diagnosis. [...]”
Invent API
“A 62-year-old male presents with progressive shortness of breath, bilateral inspiratory crackles, and a dry cough. He has a 30-year history of asbestos exposure. Discuss the differential diagnosis between idiopathic pulmonary fibrosis and asbestosis.”
“Differential Diagnosis: IPF vs. Asbestosis.
The presentation of progressive dyspnea, dry cough and bilateral ‘Velcro’ crackles with a 30-year asbestos history is a classic diagnostic challenge. Both are chronic, progressive fibrosing interstitial lung diseases sharing a Usual Interstitial Pneumonia (UIP) pattern. [...]”
表 5:各 API 為醫療問答資料集產生的提示-完成對範例,其查詢內容為:「涵蓋診斷、治療、藥物說明及醫療角色之醫療問答對資料集。每個項目包含一個提示以及一份詳細的專家生成回應。」
2.2 評估指標
多樣性指標。我們從兩個互補的軸向來衡量生成資料集的多樣性,即語義與詞彙。
- 語義多樣性。我們主要採用基於語義的 DCScore (Zhu et al., 2025) 作為衡量指標。DCScore 是一種基於嵌入的指標,將多樣性評估視為一種樣本分類任務,使該指標能捕捉各項指標之間的相互關係。
樣本。我們使用 Qwen-3-4B-embedding 模型 (Zhang et al., 2025) 作為嵌入模型來計算多樣性分數,因為該模型支援多種語言與 32K 的上下文長度,這對於正確評估基準測試套件中的資料集至關重要。
- 詞彙多樣性。除了 DCScore 之外,我們也報告與文本形式相關(而非語義內容)的指標,針對產生的提示進行計算:(i) n-gram 多樣性 (NGD),即語料庫中 n ∈ {1, ..., 4} 的唯一 n-gram 與總 n-gram 之比的總和 (Shaib et al., 2026; Padmakumar & He, 2024),數值越高代表多樣性越高;(ii) 壓縮比 (CR),即原始語料庫大小與壓縮後語料庫大小的 gzip 比率 (Shaib et al., 2026),數值越高代表冗餘度越高;以及 (iii) 完全重複率,即逐字重複提示的比例。
在所有這些指標中,我們針對不同資料集規模進行比較時,會從每個產生的語料庫中抽取固定 2,000 個提示的隨機樣本來報告結果。
品質指標。我們採用基於評分準則的 LLM 評審方法來評估品質,因為多項研究已證實此方法與人類判斷具備高度相關性(Zheng et al., 2023; Liu et al., 2023)。
- 提示與完成品質的評分採 0–10 分制。這些指標的詳細描述、完整的裁判提示(見 A.1)以及各資料集規模下的評估結果(見 A.2),均已於附錄中說明。 * 相關性(Relevance)衡量生成樣本在回應使用者請求的意圖與範疇時的契合程度,評分同樣採 0–10 分制。相關性分數可驗證樣本是否符合請求的任務類型與主題,且是否以目標語言撰寫,並遵循請求中明確指出的約束(例如:長度、格式、結構、風格)。 * 後訓練品質(Post-training quality)我們採用 LLM-as-a-Judge 排名機制,比較各生成器在醫療問答(medical QA)資料上微調後的模型表現。評估採用一組內部保留(held-out)的醫療領域提示。我們不從單一生成器的資料中劃分測試集,因為這類分割會來自該生成器自身的分布,進而對該模型產生偏見(Torralba & Efros, 2011; Teney et al., 2020)。我們的測試提示是從真實世界的醫療查詢分布中抽樣,並從所有生成器的訓練資料中保留。這是一個更具挑戰性的分布外(out-of-distribution)測試,也是下游泛化能力的更可靠指標(Recht et al., 2019; Koh et al., 2021)。針對每個測試提示,裁判(Gemini 3.1 Pro)在單次呼叫中檢視所有回應,並回傳最佳至最差的排序。為了控制位置偏見(position bias),每次提示的回應順序都會獨立隨機洗牌,確保沒有任何模型佔據固定的位置。
我們在附錄 A.8 中提供了用於排名評估的裁判提示,供讀者參考。
3 結果
3.1 整體效能:多樣性與品質
多樣性與品質比較 圖 1 繪製了在 5,000 樣本規模下,所有八個資料集平均品質與平均多樣性(DCScore)的關係。一個強大的合成資料產生器應位居右上區域,展現其兼具多樣性與高品質。Invent API 是唯一達成此目標的產生器,在兩個維度上的表現均遠超所有外部系統。它以大幅領先的優勢實現了最高的多樣性,分別比次佳模型(GLM-5.3)高出約 19%、比 Claude Opus 5 高出 24%、比 GPT-5.6 Sol 高出 29%、比 Gemini-3.1-pro 高出 32%,以及比 DeepSeek-V4-Pro 高出 55%。同時,它也獲得了最高的品質分數(≈ 7.9),比最強基線(Claude Opus 5 與 GLM-5.3,兩者均約為 6.75)高出約 17%。附錄 A.7 報告了在所有四種資料集規模下的多樣性與品質權衡。
至關重要的是,Invent API 的多樣性優勢並未以品質為代價,它位於帕累托前緣,能同時提升這兩項指標。在 5,000 樣本的規模下,平均於所有資料集上的表現(圖 6a)為 7.90,是品質最高的生成器,領先 Claude Opus 5 (6.76) 與 GLM-5.3 (6.73) 約 17%,且遠高於 Gemini-3.1-pro (5.26) 與 GPT-5.6 Sol (5.13)。
Diversity and Quality across different languages INVENT-A-DATASET supports 244 different languages (Hooker & Roy, 2026), hence it of interest to understand performance across different languages. In Figure 3 we compare Serbian and Hindi (Figure 3a), and show the Invent API is clearly the most diverse (≈ 0.082, roughly double the next-best baseline) while retaining high-quality generations; GPT-5.6 Sol reaches slightly higher quality (≈ 8.05) but at less than half the diversity, leaving the Invent API alone in the upper-right. Across six English datasets (Figure 3b), the separation is cleaner. The Invent API leads on both axes, with the highest quality (≈ 7.5) and the highest diversity (≈ 0.149, ahead of GLM-5.3 at 0.135). The multilingual setting therefore reinforces the English results, the Invent API's diversity advantage is not language-specific. Comparing the Invent API to itself, quality is nearly unchanged across languages (≈ 7.5 in English vs. ≈ 7.4 in Serbian and Hindi), while absolute diversity is lower (≈ 0.149 vs. ≈ 0.082); however, every baseline loses more diversity (roughly 60–75% vs. ≈ 44%), so the Invent API's relative advantage widens outside English.
我們也評估產出的內容是否為所要求的語言;印地語(Hindi)的語言混淆(language confusion)結果(Marchisio et al., 2024)請參考附錄 A.4。為了驗證生成資料集的語言,我們對每個生成的範例套用了 langid (Lui & Baldwin, 2012)。所有評估的基線(Baseline)皆展現極低的語言偏移(language drift),產生的大多數範例均為目標語言(target language);其中最高的偏移速率(Rate)為 7.25%,出現在 DeepSeek-V4-Pro。INVENT-A-DATASET 在所有方法中展現最低的偏移,其印地語輸出的偏移率僅 0.15%(低於 1%),被識別為其他語言的比例極低。
3.2 約束對資料集多樣性的影響
實務資料請求通常會施加明確的約束(長度、格式、結構、人設),這會縮減有效輸出空間,並通常降低多樣性。這些約束通常需要更複雜的生成模式,進而提高合成資料生成的難度。圖 4 比較了不受約束的資料集查詢與受約束查詢(參見表 3 與表 8)的多樣性。在所有外部基準中,加入約束都會降低多樣性,降幅最高達 22%。GPT-5.6 Sol 降幅為 22%(0.0328 → 0.0257)、DeepSeek-v4-Pro 為 18%(0.0264 → 0.0217)、Gemini-3.1-pro 為 17%(0.0306 → 0.0254)、Claude Opus 5 為 15%(0.0248 → 0.0210)以及 GLM-5.3 為 15%(0.0317 → 0.0271)。Invent API 是唯一多樣性不受影響的管線(兩種設定均為 0.0428),且沒有任何基準能接近其穩健性。因此,在受約束的情況下,Invent API 的多樣性為所有管線中最高(0.0428),比第二名(GLM-5.3,

圖 3:多語言設定下的多樣性與品質。在受約束設定下,針對非英語資料集抽取 5,000 個樣本時的多樣性與品質權衡。在兩種設定下,Invent API 均位於右上區域,在多樣性表現上始終領先,且在六種額外語言中的品質表現也位居前列。
0.0271)。產生符合精確約束的資料,對於現實世界的建模至關重要;這也顯示 Invent API 能在約束原本會降低多樣性的情況下,精確地維持多樣性。
3.3 規模化下的多樣性與品質
多樣性隨規模擴大而下降。我們也測量了多樣性如何隨資料集規模演進。在圖 5b 中可觀察到兩種效應。首先,隨著資料集規模擴大,所有外部生成器的絕對樣本多樣性均呈現下降趨勢:這是預料之中的結果,因為每個新樣本都必須與已生成的龐大樣本庫有所區別,因此在相同提示下,單一模型會越來越依賴重複的模式,且每增加一個樣本的邊際多樣性增益也隨之穩定縮減。第二點,且更重要的是,Invent API 的相對優勢會隨規模擴大而增強。在 2,000 個樣本時,Invent API 的平均多樣性已領先最強的基線(glm-5.3)0.037(0.250 對 0.213,相對增益為 17%);而在 20,000 個樣本時,此差距擴大至 0.072(0.268 對 0.196),相對增益達到 37%,優勢更超過一倍。這種差距擴大的原因在於趨勢分歧:Invent API 的多樣性隨資料集規模增加而提升,而所有基線模型的多樣性則要麼進入收斂高原(deepseek-v4-pro、gemini-3.1-pro),要麼呈現下降(glm-5.3、gpt-5.6-sol、claude-opus-5),這顯示前沿 LLM 隨著生成資料量增加,會越來越頻繁地重複自身內容。
品質與規模。如圖 5a 所示,Invent API 的品質領先優勢在不同規模下皆保持穩定。相較於 GLM-5.3,Invent API 在各種規模下均具備領先地位(200 規模時為 8.43 vs. 7.03,2,000 規模時為 8.09 vs. 6.76,5,000 規模時為 7.90 vs. 6.73,20,000 規模時為 8.15 vs. 6.55),整體相對優勢約維持在 17–24% 左右。隨著生成規模擴大,Invent API 的分數維持在一個狹窄的區間內(7.90–8.43),而 GLM-5.3 則持續下滑(7.03→6.55),這顯示 Invent API 的品質即使在資料集規模增加且多樣性壓力提升的情況下,依然能保持穩定。結合多樣性指標的結果,這證明了多樣性與品質在擴展規模時皆能持續存在。
規模與詞彙多樣性。如第 2.2 節所述,我們在基於語義內容的 DCScore 之外,額外採用了三個基於文本形式的多樣性指標。表 6 報告了

圖 4:根據生成器,在受約束與不受約束設定下的多樣性。以 DCScore($\tau = 0.1$)衡量,並針對每個包含 20,000 個樣本的 8 個資料集進行平均。Invent API 的表現最為穩健:其多樣性基本上沒有變化(兩種設定均為 0.0428),而每個基線(Baseline)則下降了 15-22%。因此,在受約束的情況下,Invent API 的多樣性為所有管線(Pipeline)中最高(0.0428,而次佳的 GLM-5.3 為 0.0271)。
在 2,000、5,000 與 20,000 的資料集規模下,於不受約束(unconstrained)情境中評估,且每個規模皆以固定的 2,000 則隨機提示(prompt)樣本進行計算,使各規模間的數值具備直接可比性。Invent API 在所有指標與所有規模下皆為表現最佳的生成器。其 n-gram 多樣性(Diversity)為最高,且隨資料集規模增加而提升(1.75 → 1.82 → 1.87)。相較於最接近的基線(GLM-5.3),其在 2K、5K 與 20K 規模下的相對增益分別為 9.4%、7.1% 與 10.7%。此外,在所有規模下,Invent API 均呈現最低的壓縮比(≈ 3.5,即其提示重複性最低;相較之下,基線的數值為 3.95-5.5)。其在所有規模下的完全重複率(exact-duplicate rate)均為零(0.0%),而所有外部基線的完全重複提示比例則為 6.2-19.2%,且隨資料集規模增加而上升,其中以 DeepSeek-V4-Pro 的表現最差。這些指標顯示,Invent API 的提示最為多樣、最難壓縮且不具重複性,而前沿 LLM 則隨著生成資料量增加,越來越傾向於重複。受約束(constrained)情境的結果則詳見附錄 A.5。
3. 4 生成樣本的相關性
生成樣本的相關性 相關性衡量樣本是否符合資料集規格,與其建構品質或多樣性無關。圖 6b 顯示在 5,000 個樣本中,各模型的平均相關性。雖然所有生成器的相關性均相當高(7.1–8.9),證實每個生成器皆能產生符合規格的資料,但 Invent API 的表現位居中段而非頂尖。Claude Opus 5 (8.86) 與 GLM-5.3 (8.66) 的表現比 Invent API 高出一個多百分點。Invent API (7.47) 的表現與 Gemini-3.1-pro (7.51)、DeepSeek-v4-pro (7.44) 以及 GPT-5.6 Sol (7.11) 相當。Invent API 目前的優勢在於維持多樣性與品質,但目前並非總是能產生最符合規格的相關樣本。在不犧牲多樣性的前提下提升相關性,仍是 Invent API 未來重要的改進方向。

圖 5:左:品質與規模。GLM-5.3 與 Invent API 在 200、2,000、5,000 與 20,000 樣本下的平均資料集品質。Invent API 在各個規模下均領先,且隨著資料集擴大仍保持穩定。右:多樣性與規模。在 2,000、5,000 與 20,000 樣本下,所有生成器的平均多樣性(DCScore,$\tau = 0.1$),係以從各資料集中隨機抽取 2,000 個樣本進行測量。對於大多數基線生成器而言,樣本規模越大,每樣本的多樣性就越低;相較之下,Invent API 的多樣性仍維持在高點(甚至有所提升),因此其相對優勢會隨著規模擴大而增加。

圖 6:左:在 5,000 樣本規模下,資料集提示與完成品質的平均值。Invent API 是品質最高的資料集產生器,領先於 Claude Opus 5 與 GLM-5.3。右:各產生器的相關性。在 5,000 樣本規模下的平均樣本相關性。所有管線產出的資料皆符合規格(7.1–8.9);Claude Opus 5 與 GLM-5.3 領先,而 Invent API 則具備競爭力。
3.5 對下游後訓練品質的影響
為了測試生成資料品質是否能提升模型表現,我們針對三款模型重複以下流程:Llama-3.3-70B、Gemma-4-31B-it 與 Qwen3.5-9B。針對每一款模型,我們將未經微調的基底模型與五個獨立微調版本進行比較,這些版本皆是在使用不同生成器產出的 20K 筆醫療問答資料集上進行訓練。所有六款模型皆須回答同一組保留的醫療提示(medical prompts)。我們使用 LLM 評審(LLM Judge)在單次呼叫中,將這六個回應從最佳到最差進行排序。
在所有微調的模型中,我們觀察到使用 Invent API 資料訓練的模型表現明顯優於其他模型。如圖 7(左)所示,Llama-3.3-70B 在使用 Invent API 資料訓練的模型中,有 54% 的樣本排名第一。這比基礎未微調模型(25%)高出一倍以上,也比在 Claude Opus 5 生成資料上微調的最強競爭模型(9%)高出六倍。這些模式同樣適用於 Gemma-4-31B-it(圖 7,右),在其他生成器資料上微調的四個模型均未能超越基礎模型在 Top-1 排名佔比上的表現(DeepSeek-V4-Pro (12%)、Claude
1.75 1.82 1.87 3.67 3.51 3.47 0.0 0.0 0.0 1.57 1.70 1.69 4.70 3.95 3.96 6.2 6.8 8.9 1.53 1.62 1.66 4.59 4.22 4.08 7.5 6.8 6.5 1.60 1.65 1.59 4.33 4.33 4.39 9.6 10.4 12.4 1.55 1.54 1.48 4.27 4.25 4.28 8.5 7.9 8.4 1.07 1.07 1.07 5.50 5.53 5.51 19.2 18.5 18.8
表 6:在不受約束的情境下,八個資料集平均的詞彙提示多樣性與重複率。每個資料集大小(2K/5K/20K)均以固定的 2,000 個提示隨機樣本進行測量,因此各欄位可直接進行比較。N-gram 多樣性 (NGD):↑ 代表多樣性更高。壓縮比 (CR):↓ 代表冗餘度較低;Dup 為完全重複提示的百分比(↓ 代表表現較佳)。各欄位中表現最佳的資料變體以粗體標示。
Opus 5 (10%)、GLM-5.3 (8%) 與 GPT-5.6 Sol (8%)。至於 Qwen3.5-9B,其未經微調的基礎模型是一個非常強大的起點,在 Top-1 排名佔比上略勝 Invent API 微調模型 (44% vs. 38%),然而 Invent API 微調模型仍排名在其他所有生成模型的微調版本之上 (詳情請參閱附錄 A.6)。這些結果共同證實,資料品質與多樣性的差異會直接影響在該資料上訓練出的模型表現。此排名順序與第 3.1 節及 3.3 節的資料集層級結果一致。
4 相關研究
Historically, significant effort has been invested in collecting, labeling, filtering, and reshaping data to fit a training objective (Wang et al., 2024a) A large body of work has emerged which shows that efforts to better curate training corpus, including de-duping (Taylor et al., 2022; Kocetkov et al., 2022), data pruning (Marion et al., 2023; Singh et al., 2024b; Sorscher et al., 2022; Albalak et al., 2024; Tirumala et al., 2023; Chimoto et al., 2024) or data prioritization (Boubdir et al., 2023a; Thakkar et al., 2023) can lead to better model quality. However, most of these works start from an existing data pool and aim to modify it to arrive at higher quality data. The result is often a dataset that only approximates the desired behavior, and a model trained around the limitations of the available data rather than the target capability. In contrast, our work is focused explicitly on the zero data regime where no data is available and optimizing directly for any set of constraints or capabilities a practitioner desires.
Dataset search approaches Dataset search poses unique challenges distinct from traditional web search (Brickley et al., 2019). Users span a wide range of expertise and goals, and in many cases those goals are far removed from the datasets themselves. The datasets are also often difficult to peruse manually. Despite advances in interpreting natural-language intents (Chapman et al., 2020a), users still struggle with incomplete and inconsistent metadata, with expressing information-seeking needs as structured search constraints, and with assessing dataset relevance (Kacprzak et al., 2019; Li et al., 2025b). Crucially, dataset search is fundamentally limited to retrieving data that already exists; when no suitable dataset is available, search offers no recourse.
Synthetic Data Synthetic data offers several compelling advantages. It allows practitioners to

圖 7:下游排名評估中的 Top-1 排名佔比。圖中顯示各個微調模型在保留的醫療測試提示中排名第一的百分比,所有六個回應皆在單次呼叫中由裁判共同評分。每個面板的 x 軸列出了在不同合成資料變體上微調的模型,以及未經微調的基礎模型(灰色)作為對照組。左圖:Llama-3.3-70B。Invent API 微調模型表現最為突出,在 54% 的提示中排名第一,是未經微調基礎模型(25%)的兩倍多。右圖:Gemma-4-31B-it。Invent API 微調模型再次領先,排名第一的比例為 41%,遠高於基礎模型(21%)。在兩種情況下,使用其他生成器(GLM-5.3、Claude Opus 5、GPT-5.6 Sol、DeepSeek-V4-Pro)資料微調的模型,其表現均遠遜於基礎模型。
針對真實分布中罕見的長尾情境(Long-tail Scenarios),我們的方法能降低額外資料收集的成本,並能透過自我遞迴的方式逐步提高範例難度,藉此激發出更進階的能力(Xu et al., 2024; Luo et al., 2023)。隨著時間推移,生成新資料的成本已大幅降低,進而實現了資料空間中的「即時」(on-the-fly)最佳化(Shimabucoro et al., 2024)。然而,這些優勢取決於生成資料的多樣性。若在狹隘或重複的合成樣本上進行訓練,將導致效能下降,且在遞迴極限下會引發「模型崩潰」(Model Collapse),使後續模型逐漸失去分布的長尾部分(Shumailov et al., 2024; Schaffelder & Gatt, 2026)。因此,維持資料多樣性是任何合成資料生成方法的核心目標。多數資料合成方法會利用現有的資料結構,並將其調整以符合新需求。以 Self-Instruct(Wang et al., 2023b)和 Alpaca(Taori et al., 2023)等基於種子(Seed)的方法為例,這些方法能從少量的人工撰寫種子範例中,引導(Bootstrap)出大型指令微調(Instruction-tuning)資料集;而 Evol-Instruct/WizardLM(Xu et al., 2024)及其程式碼變體 WizardCoder(Luo et al., 2023)等演化式方法,則是透過迭代重寫現有資料集,將其轉化為更複雜且多樣的實例(Instance)。其他研究則會針對特定目標模型或任務量身打造合成資料,但仍需依賴種子指令或現有語料庫(Corpus)作為條件(Wang et al., 2024b)。相較之下,本研究假設最終使用者完全無法存取任何現有資料。
我們僅根據對目標任務的自然語言描述,建構出一套零樣本資料集。
5 結論
直接使用外部模型 API 的一個系統性弱點在於,其產生的樣本容易陷入重複的模式,進而削弱產出資料集的多元性。當追蹤資料集大小與多元性之間的關係時,此效應尤為明顯。當對單一模型發出重複提示時,產生新穎輸出的難度會隨之增加,且每筆樣本的多元性也會持續下降。相較之下,Invent API 對此類「塌陷」現象的抵抗力遠為強大,且隨著資料集規模擴大,其優勢便會顯現。在擴大規模後,Invent API 產出的資料不僅比所有基線模型更具多元性,同時也能維持最高的整體資料集品質,並確保結果符合規格。
限制 ---
INVENT-A-DATASET 存在一些重要的缺點。目前該工具主要用於產生 SFT 與對齊訓練的指令與偏好資料集,但尚未擴展至模型在驅動 Harness 時所產生的工具呼叫軌跡或多步驟、智能體式的軌跡。我們將此留待未來研究,因為 Harness 資料集通常與特定環境緊密相關:包含可用的工具、工具的簽章(signatures)以及 Harness 的控制流。這種耦合使得這類資料集的適用範圍較窄,且隨著環境演進,資料集容易過時,相較於目前支援的常規 SFT 與偏好資料集,其跨情境的重複使用難度也較高。此外,另一個值得探討的問題是,這類資料是否最適合透過參數權重更新來內化。再者,INVENT-A-DATASET 目前僅限於文字模態,未來我們希望能擴展以支援多模態資料集生成。
附錄
A.1 品質分數描述
提示品質。本評分準則根據具體性、範疇與完整性等標準,以 0–10 的量表評分。低分(1–4 分)代表內容不連貫、模糊或規格不足;中分(5–6 分)代表清晰且可執行的任務,但缺乏深度或上下文;高分(7–10 分)則為結構良好且具備明確成功標準的提示。最重要的是,提示是在資料集要求的格式內進行評估,因此若使用者要求提示簡潔,則不會因其簡短而扣分。
完成品質:根據生成回應的整體品質、結構以及對提示語句所列約束的遵循程度,以 0–10 分制進行評分。低分(1–2 分)代表內容不連貫、錯誤或不完整;中等分數(3–5 分)則代表內容專業且準確,涵蓋了主要重點但缺乏深度;高分(6–10 分)則代表內容詳盡且組織良好,能處理疑慮與邊緣案例。與提示品質相同,深度的評估是相對於資料集請求所設定的長度、格式或風格限制。
以下為完整的裁判提示模板:
3 品質與相關性裁判模板
嚴謹且細緻的裁判,用於彙整監督式微調資料集。給定目標規格(包含資料集層級的資料請求與語言標籤)以及一組候選提示/完成對,裁判會決定該候選是否應納入資料集,以及其品質評分。
- 評分維度:相關性、提示品質、完成品質——各項評分為 0–10 的整數。
- 輸出:單一精簡的 JSON 物件,不含其他文字。
- 模板佔位符:
{query}、{language}、{prompt}、{completion}、{prompt质量_rubric}、{completion质量_rubric}。
3. System Message
你是一位嚴謹且細心的裁判,負責彙整監督式微調資料集。根據給定的目標規格(資料集層級的資料請求與語言標籤)以及一組候選提示/完成對,你需決定該候選內容是否確實屬於該資料集,並評估其品質。請保持果斷且一致:針對準確、完整且主題相關的提示,並搭配精確、完整的完成內容給予獎勵;對於主題不符、語言錯誤、
不完整/被截斷、答案式或元指令(例如「產生一個資料集...」)的提示,以及內容空洞、錯誤或閃爍其詞的完成內容則予以扣分。你會根據三個獨立的評分維度為候選內容評分——相關性(該樣本是否屬於資料集?)、提示品質(該提示本身的品質如何?)以及完成品質(該完成內容作為提示的答案品質如何?)——每個維度的評分範圍為 0-10。請僅回覆一個簡潔的 JSON 物件,且不得包含任何其他文字。
提示品質評分準則
為提示的整體品質評分,範圍為 0 到 10 分。
規模
0 (空提示) — 沒有提示。這是一個空值項目。
1-2 (不連貫 / 無意義) — 請求內容無法閱讀、隨機且缺乏意義。僅有單字。胡言亂語、不相關的片段,且沒有明確的任務。根據現有的上下文,可能無法回答。無法辨識出任何任務、目標或意義。範例:「asdf quantum banana」、「Help」、「AI」、「Write something」、「????」
3-4 (不明確或規格不足) — 內容可讀,但缺乏足夠細節以有效執行。雖然可以辨識出主題,但缺少了指定範疇、受眾、格式或約束等關鍵參數。請求過於籠統。範例:「告訴我關於科技的資訊」、「撰寫關於氣候變遷的文章」、「解釋機器學習」、「你覺得如何?」(【譯註:此處原文為「What do you think?」,意指引導受試者表達個人觀點的開放式問題。】)
5-6(基本清晰度)— 具備可識別的任務與明確的動作動詞。內容直觀但缺乏深度或策略性上下文。範例:「摘要這篇文章」、「為初學者撰寫 3 段落的說明」、「列出遠端工作的 5 項優缺點」。
7-8(結構良好)— 包含相關上下文或使用案例。具備明確的成功標準或交付規格。可能要求進行比較分析或結構化推理。範例:「為科技電子報受眾撰寫一篇 150 字的 AI 安全論文摘要,並強調實務影響」。
9-10(專家級/元最佳化)— 運用進階提示工程與精密的鷹架設計。明確指派角色或人設規格。具備多維度的推理或評估標準。預先考量邊緣案例、細微差別或品質標準。可能包含自我修正機制或迭代式指令。
特別案例
- 尊重資料集請求的約束:請勿因提示缺乏細節、長度、上下文或鷹架而扣分,前提是這些內容是資料集請求明確要求省略或限制的。當請求規定了特定格式(例如:短提示或單行提示、不提供上下文、固定模板或口語/非正式語體)時,請評估該提示在該格式內是否精煉:其內容是否清晰、具備目的感,且符合其預期範疇?一個刻意簡潔且明確規範任務的提示仍可獲得 7 分以上;請將 3-4 分保留給那些在請求要求之外顯得模糊或規格不足的提示。請根據請求的格式評估規格品質,而非以不受約束的理想標準進行評比。
- 多輪對話:若提示涉及多輪對話,請根據對話中的最後一個回合進行評分。
- 複雜且表述不佳的提示:請專注於提示的結構與規格。
複雜度並非重點。即便問題困難,只要提問方式不佳或預期產出不明確,品質評分仍應偏低。
- 長提示不代表高品質:請留意那些看似詳盡、實則缺乏重點的冗長提示。若請求內容過於繁瑣且缺乏有助的結構,應將其評為低品質。
完成品質評分準則
針對完成內容,請根據提示(Prompt)為其評分,評分範圍為 0 到 10 分。
規模
0(空回應)— 無回應;為空值。
1(不連貫 / 無意義)— 回應無法閱讀,答案完全錯誤,或未能回應提示。內容為胡言亂語、斷句錯誤或語意空洞的輸出。完全誤解或忽略提示。討論完全錯誤的主題或提供無關資訊。包含關鍵事實錯誤,導致回應無法使用。可能僅提及提示,但未提供實際答案。
2(不足或不完整)— 回應提示,但存在重大缺漏或問題。僅提供極其基本或表面化的資訊。回答提示的部分內容,卻忽略其他部分。存在顯著的準確度問題或誤導性陳述。缺少關鍵上下文、細節或資格說明。過度簡化,以致於毫無幫助。內容雜亂無章或結構不佳。
3–5(基本清晰度)— 具備能力且符合基本要求的回應。直接回應提示。內容事實正確且無重大錯誤。對主要觀點的覆蓋率合理。結構清晰且易於理解。可能包含基本的範例或上下文。足以供基本理解,但缺乏深度或細微差別。
6–8(結構良好)— 內容涵蓋全面,並包含相關細節與上下文。具備深刻且細緻的理解,並能從多個角度切入。包含有助益的範例、類比或插圖。組織結構良好,邏輯流程清晰。探討了邊緣案例、注意事項、限制或權衡。能預見後續問題或潛在的誤解。提供實用且可執行的指導。
9–10 (專家級 / Meta-Optimized) — 實質超越所要求的內容(以有幫助的方式)。展現具備策略深度與綜合能力的精通。提供新穎見解、非顯而易見的連結或可轉移的架構。識別未明說的假設或處理潛在需求。結構設計以追求極致的清晰度與可用性(可能包含表格、Schema、決策樹)。預測並防止常見錯誤或誤解。教導底層原理,而不僅僅是回答提示。
特別案例
尊重資料集請求的約束:請勿因完成品缺乏深度、長度、詳述、範例或細微差別而扣分,前提是這些項目在資料集請求中已明確排除或設有上限。當請求明確規範完成品的長度、格式、風格或深度(例如:「僅限單句回答」、「簡潔」、「僅限項目符號列表」、「不需解釋」),請評估該回答在這些限制範圍內,是否能完整且準確地滿足提示。只要回答簡潔、正確且執行得當,並符合簡短或格式要求,仍可獲得 6 分(或更高)的評分。
它能完美執行受約束的形式(例如,執行得非常出色);請勿將要求的簡潔或簡明視為「不充分」或「流於表面」。請相對於要求的形式來評估深度,而非與不受約束的理想狀態進行比較。
- 多輪對話:將完成內容視為對最後一輪使用者回合的回應進行評分。
- 長篇完成內容不一定代表高品質:冗長、填充或重複且未增加實質內容的回答,應被評為低品質。
完整裁判提示
目標規格
- 資料集資料請求(任務與主題;適用於整個資料集):
{query} - 目標語言:
{language}
針對包含提示與完成內容的候選樣本,請在三個獨立的維度上進行評分,每個維度的評分為 0-10 的整數。
軸 1 — relevance_score
候選 SAMPLE 是否屬於目標資料集?歸屬資格需符合四大硬性標準,且所有標準皆須達成。這些標準依重要性排序如下:(A) 與 (B) 為基礎,而 (C) 的優先級高於 (D):明確違反明確陳述的限制,其歸屬失敗的嚴重程度高於提示可用性缺陷。
(A) 任務與主題:樣本須符合資料集要求的任務類型與主題。 (B) 語言:樣本須使用資料集要求指定的目標語言撰寫。 (C) 遵守約束:樣本須遵守資料集要求針對提示與/或完成內容 EXPLICITLY 說明的約束(例如:長度、格式、結構、風格、人設)。僅計算要求實際說明的約束;若要求未說明任何約束,則視 (C) 為已滿足,請勿自行發明約束。 (D) 可用提示:提示須為自給自足、完整且獨立的使用者提示,而非答案/完成內容、非截斷片段、非關於建立或生成資料集的後設指令、非離題填充內容,且不應依賴幻覺或缺失的上下文(例如:提及實際上未包含在提示中的檔案、文件、圖像、表格、連結或先前對話)。
3 relevance_score scale
0(空)— 無提示或完成內容 / null / 無法讀取。 1–2(不屬於)— 失敗(A)或(B):主題與資料集請求的任務或主題無關、任務類型錯誤,或語言錯誤。例如:目標為「客戶支援詢問」,候選內容為「粒線體是細胞的發電廠」。 3(違反限制)— 主題相關且為該語言,但明顯違反明確陳述的(C)限制(例如:格式錯誤、超出指定長度、缺少必要結構)。由於(C)優先於(D),此項評分應低於單純的可用性缺陷。
4(無法使用)— 符合主題、使用目標語言,且未明顯違反所述約束,但仍不符合 (D) 類別:其內容為答案或完成品而非提示,或是僅為片段、無意義的文字、被截斷或不完整的請求、元指令或資料生成指令(例如「生成 50 張帳單爭議票」),或依賴不存在或虛構的上下文(例如「摘要附加檔案」、「根據上方表格」)。5–6(邊緣成員)— 符合 (A)、(B)、(D) 且未出現明顯的 (C) 違規,但僅部分或鬆散地遵循所述約束,且/或內容過於籠統、規格不明確,或僅與請求的主題相關。7–8(符合)— 符合 (A)、(B)、(D) 並完全遵循所有所述約束 (C):明確符合任務、主題與語言,滿足所有請求的約束,且為一個完整、具體、獨立的提示,具備足夠的具體性以作為優質範例。僅有微小缺漏。9–10(極佳符合)— 符合所有四項標準,完全且精確地遵循所有所述約束,且內容具備真實感、具體性,並完全自給自足——為此資料集中的理想成員。相關性說明。首先閱讀自由格式的資料集請求,以擷取 (A) 任務與主題、(B) 任何語言要求,以及 (C) 任何明確的格式、長度、格式、風格或人設約束。優先順序為 (A)/(B) > (C) > (D):語言或主題錯誤的上限為 1–2;明確違反所述約束 (C) 的上限為 ≤ 3;可用性缺陷 (D) 的上限為 ≤ 4。若要得分 7+,必須完全符合所有所述約束;部分或鬆散符合的上限則為 5–6。
若請求中未設定任何約束,則須符合條件 (C) —— 不得限制其分數,亦不得自行添加約束。針對多回合候選項目,請以使用者的請求為錨點,評估整段對話。軸向 2 — 提示品質_score
在不考慮目標資料集的情況下,該候選提示本身的品質如何?請參考上述「提示品質評分準則」({prompt_quality_rubric});僅針對使用者提示進行評分。提示品質與相關性(Relevance)無關:離題的提示仍可能寫得很好(高品質、低相關性),而高度相關的提示也可能粗糙或規格不足(高相關性、低品質)。請分別為這兩個軸向評分。然而,若資料集請求明確要求特定屬性,則不應因該屬性而扣減提示品質。若請求要求特定格式(例如:短提示/單行提示、不提供上下文、固定模板、純口語化措辭、特定長度或格式),請評估該提示在該格式下的執行成效,而非評估其深度、長度或是否具備「鷹架」(Scaffolding)等在不受約束情況下才能獲得高分的特質。若提示因請求要求而保持簡潔,則不應被視為「規格不足」。
軸向 3 — 完成品質_score
該候選完成(Completion)作為對提示的回應品質如何?請參考上述「完成品質評分準則」({completion_quality_rubric});僅針對完成內容進行評分,並以提示的要求與資料集請求的預期為依據。
完成品質與另外兩個維度無關:對於離題或粗糙的提示,即使給出極佳的回答,在此維度中仍會獲得高分;而對於優質提示給出薄弱的回答,則會得到低分。然而,若資料集請求明確要求了某項屬性,請勿因該屬性而扣減完成品質的分數。若請求規定了完成品的長度、格式、風格或深度(例如:「單句回答」、「簡潔」、「僅限項目符號清單」、「不需解釋」),請評估完成品的表現如何。
在這些限制範圍內,重點在於是否能滿足提示要求,而非該答案是否具備更豐富的闡述、範例或細微差別,以在不受約束的回答中獲得更高分。
### 決策
僅當 (A) 與 (B) 與 (C) 與 (D) 全部成立時,才將 relevant 設為「yes」;這通常對應 relevance_score ≥ 5;否則,將其設為「no」(relevance_score ≤ 4)。(C) 當樣本未明顯違反任何明確陳述的限制時,即為「成立」;(D) 當提示符合上述定義時,即為「成立」。請確保 relevance_score 與 yes/no 決策保持一致。
3. 僅回傳此 JSON 物件,且不得包含其他內容
{"relevance_score": <int 0-10>, "prompt_quality_score": <int 0-10>, "completion质量_score": <int 0-10>, "relevant": <"yes" or "no">, "reason": <("string" ≤ 30 words naming the deciding criteria for relevance, prompt quality and completion quality>)}
候選提示:{prompt}
候選完成內容:{completion}
### A.2 各資料集的品質指標
圖 8 與圖 9 分別呈現了在所有 4 個尺度下,受約束與不受約束設定的各資料集品質結果。
### A.3 資料集生成查詢
表 8 列出了用於建立評估集內其他資料集的受約束與不受約束查詢。
### A.4 輸出語言一致性
### A.5 詞彙多樣性:受約束情境
表 11 顯示在受約束情境下,相同的三個詞彙提示指標於 2,000、5,000 與 20,000 樣本規模(各基於固定的 2,000 個隨機提示樣本)的表現。其模式與表 6 中的不受約束結果相似。Invent API 具有最少冗餘提示(在所有規模下皆具備最低壓縮比),且在所有規模下的重複率均接近於零;相較之下,基線的重複率則隨資料集規模而增加。在原始 *n*- gram 多樣性方面,Invent API 與 GLM-5.3 表現相近且交替領先(在 5,000 樣本時,GLM-5.3 略微領先)。整體而言,受約束的生成...
<table><thead><tr><th>Dataset</th><th>Unconstrained Query</th><th>Constrained Query</th></tr></thead><tbody><tr><td>Customer Support Banking</td><td>Dataset of customer service responses guiding users through credit card activation, blocking, and mortgage inquiries.</td><td>Dataset of customer service conversations guiding frustrated users through credit card activation, blocking, and mortgage inquiries. The customer service agent is always welcoming and polite.</td></tr><tr><td>Hindi News</td><td>Dataset containing Hindi news articles sourced from various Indian news websites, paired with their corresponding headlines and summaries.</td><td>Dataset containing Hindi news articles sourced from various Indian news websites about sports, paired with their corresponding summaries. The prompt is always the article headline, and the completion is the article summary.</td></tr><tr><td>Customer Support General</td><td>A dataset of customer support inquiries and corresponding responses, covering issues such as product information, technical troubleshooting, billing disputes, and service returns.</td><td></td></tr><tr><td>Medical QA</td><td>Dataset of medical question-answer pairs covering diagnoses, treatments, drug explanations, and healthcare roles. Each entry includes a prompt and a detailed, expert-generated response.</td><td></td></tr><tr><td>Hotel Reviews</td><td>A collection of customer reviews describing hotel stays, focusing on service, room quality, location, and amenities. The dataset should include both positive and negative experiences.</td><td></td></tr>
</tbody></table>---
表 8:用於產生資料集的單一請求查詢。選擇這些查詢是為了涵蓋 1) 任務類型、2) 語言以及 3) 領域的多元範疇。我們針對不受約束與受約束版本的查詢進行評估,其中受約束的查詢需滿足額外要求,例如 1) 語調約束、2) JSON 格式、3) 多選輸出限制。
相較於所有基線生成器,此方法能降低絕對重複率;而 Invent API 則從無約束設定下的 0.0% 略微上升至受約束設定下的 0.1%。
### A.6 後訓練結果
除了 Llama-3.3-70B 與 Gemma-4-31B-it,我們也針對 Qwen3.5-9B 進行微調。如圖 10 所示,Qwen3.5-9B 是一個非常強大的起點,在首位速率(44% vs. 38%)與平均排名(2.17 vs. 2.28)上均略勝 Invent API 微調版本。儘管如此,Invent API 微調版本仍領先於其他生成器對應的所有微調版本。
在所有選定的三種模型架構中,相對於其他生成器的微調版本,Invent API 一致地是排名最高的生成器。我們注意到,這些結果是使用僅 20K 微調資料集所取得的。先前研究顯示,監督式微調的效能會隨著訓練資料量的增加而呈現可預測的規模化趨勢(Zhang et al., 2024),因此我們預期在資料預算增加的情況下,Invent API 微調版本與最強基底模型之間的差距將會進一步擴大。
### A.7 不同規模下資料集多樣性與資料集品質的比較
圖 11–14 繪製了樣本數為 200、2,000、5,000 與 20,000 的資料集之平均多樣性與平均品質關係圖:在所有規模下,invent-api 均達成最高品質;除了樣本數為 200 時,所有 API 的表現皆相當外,invent-api 在其他規模下也具備最高的多樣性。
<table><thead><tr><th>Model</th><th>Not in language</th><th>%</th><th>Detected languages</th></tr></thead><tbody><tr><td>claude-opus-5</td><td>7/2000</td><td>0.35</td><td>hi 1993, mr 7</td></tr><tr><td>deepseek-v4-pro-0813</td><td>145/2000</td><td>7.25</td><td>hi 1855, mr 144, ne 1</td></tr><tr><td>gemini-3.1-pro</td><td>77/2000</td><td>3.85</td><td>hi 1923, mr 75, ne 2</td></tr><tr><td>gpt-5.6-sol</td><td>50/2000</td><td>2.50</td><td>hi 1950, mr 50</td></tr><tr><td>glm-5.3</td><td>49/2000</td><td>2.45</td><td>hi 1951, en 38, mr 11</td></tr><tr><td><strong>invent-api</strong></td><td><strong>3/2000</strong></td><td><strong>0.15</strong></td><td><strong>hi 1997, mr 2, lt 1</strong></td></tr></tbody></table>
<strong>1.69</strong>
<strong>1.69</strong>
<strong>1.72</strong>
<strong>3.55</strong>
<strong>3.53</strong>
<strong>3.50</strong>
<strong>0.1</strong>
<strong>0.1</strong>
<strong>0.1</strong>
1.65
<strong>1.80</strong>
1.71
4.59
3.75
3.83
0.9
1.2
3.9
1.34
1.39
1.53
5.18
4.75
4.03
2.2
2.3
1.8
1.56
1.57
1.53
4.16
4.22
4.29
2.2
3.7
5.0
1.61
1.56
1.46
4.00
4.02
4.09
2.2
3.9
4.3
1.12
1.13
1.12
5.15
5.12
5.17
11.5
11.0
11.6
表 11:在受約束設定下,八個資料集的詞彙多樣性與重複率平均值(針對每個資料集大小,以固定的 2,000 個隨機提示樣本進行計算)。N-gram 多樣性、NGD(↑ 數值越高越好),壓縮比、CR(↓ 數值越低越好),重複率(Dup %,↓ 數值越低越好)。各欄位中表現最佳的數值以粗體標示。

圖 8:針對各個生成器,在不同資料集大小 $N \in \{200, 2000, 5000, 20000\}$ 下,各受約束資料集的提示品質、完成品質與資料集相關性分數。

圖 9:針對不受約束資料集,各生成器在不同資料集大小 $N \in \{200, 2000, 5000, 20000\}$ 下的提示品質、完成品質與資料集相關性分數。

圖 10:Qwen3.5-9B 微調模型的 Top-1 排名佔比。各模型在保留的醫療提示中排名第一的百分比(每個提示會將六個回應一起進行評分);灰色代表未經微調的基礎模型。Qwen3.5-9B 的基礎模型表現異常強大,在第一名比率(44% 對 38%)與平均排名(2.17 對 2.28)上均略勝 Invent API 微調模型。儘管如此,Invent API 微調版本仍遠遠領先於其他生成器對應的所有變體。

圖 11:200 樣本資料集的多樣性與品質。每個點顯示在無約束設定下,8 個資料集中的平均品質分數(x 軸)與平均多樣性分數(y 軸,以 DCScore ($τ = 0.1$) 衡量)。

圖 12:在 2,000 筆樣本資料集中,多樣性與品質的對比。每個點代表在無約束設定下,8 個資料集中的平均品質分數(x 軸)與平均多樣性分數(以 DCScore ($τ = 0.1$) 衡量,y 軸)。

圖 13:5,000 筆樣本資料集的 Diversity 與品質。每個點代表在無約束設定下,8 個資料集平均品質分數(x 軸)與平均 Diversity 分數(y 軸,以 DCScore ($\tau = 0.1$) 衡量)。

圖 14:20,000 樣本資料集的多樣性與品質。每個點代表在無約束設定下,8 個資料集中的平均品質分數(x 軸)與平均多樣性分數(y 軸,以 DCScore ($\tau = 0.1$) 衡量)。
### A.8 排名評估
此評分準則直接比較模型,而非逐一評分。在生成的資料集上微調的模型與未微調的基底模型,皆針對相同的保留集醫療問題進行回答。針對每個問題,裁判在單次呼叫中會看到所有回應,並回傳一個從最佳到最差的排序。

這些額外細節是否真的提升了實用性,還是只是在增加冗餘內容?
**罰則:** 相較於彼此,不必要的冗長(會產生雜訊並降低清晰度)與過度簡略(遺漏使用者合理需要的上下文)都應受到懲罰。
**清晰度與自然性(低權重 — 用於打破平手僵局):**
* 回應是否語法正確且易於理解?
* 內容是否讀起來自然流暢,沒有不通順的語句或重複冗贅?
* 僅在優先順序較高的標準在各個回應之間得分相同時,才使用此標準來打破僵局。
請扮演一位公正的裁判。不要僅因回應的長度或風格而偏袒任何一方。若所有回應在實質上都無法區分,則請依最低索引值(lowest index first)進行排名。
## 輸出格式
您的輸出應為一個有效的 JSON 物件(不使用程式碼區塊),並以 UTF-8 編碼,且必須完全符合以下 JSON Schema:
```json
{
"additionalProperties": false,
"properties": {
"ranking_explanation": {
"description": "A brief rationale comparing the responses\n and explaining the ranking order.",
"title": "Ranking Explanation",
"type": "string"
},
"ranking": {
"description": "List of zero-indexed response indices\n ordered from best to worst.",
"items": { "type": "integer" },
"title": "Ranking",
"type": "array"
}
},
"required": ["ranking_explanation", "ranking"],
"title": "RankingExtractionOutput",
"type": "object"
}
您的輸出應僅包含 JSON 物件本身,不得包含任何其他文字或說明。
A.9 選定資料集樣本
下表包含每個資料集由各個生成器產出的提示與完成內容,每個資料集各取一個樣本。
| Dataset | Prompt | Completion |
|---|---|---|
| Model: Claude Opus 5 | ||
| General Customer Support | "How many people does the 8-inch cake serve?" | "Our 8-inch round cakes serve 10-12 people with standard party-sized slices, or up to 16 if you're cutting smaller dessert portions alongside other sweets. If you're expecting more than 15 guests as the main dessert, I'd suggest the 10-inch (serves 16-20) or two 8-inch cakes in different flavours, which is popular for b[...]" |
| Hotel Reviews | "Write a five-star review of a hotel that felt like great value for money." | "Five stars purely on value. Ninety-four dollars a night in a major city for a clean, quiet room with a good bed, a real breakfast included, and staff who acted like they were glad I was there. No resort fee, no parking gouge, no surprise charges. The furniture is a decade old and the lobby is plain, and I could not car[...]" |
| Customer Service Banking | "Nobody told me I needed a PIN to activate. This whole thing has wasted my morning." | "I'm very sorry that wasn't made clear to you, and I appreciate you taking the time to explain what happened. Let's make the rest of this quick. I can activate the card and set up a temporary PIN with you right now, which you can change at any ATM. Could you confirm the last four digits of your card and your mailing ZIP[...]" |
| Dataset | Prompt | Completion |
|---|---|---|
| Serbian Car Ads | “VW Caddy 1.6 TDI, 2012. godište, teretni, 301.000 km, korišćen za dostavu. Redovno servisiran na svakih 15.000 km, nedavno zamenjena turbina i kvačilo. Ima klimu, ABS, servo volan, radio i pregradnu mrežu.” | "{"vehicle_description": "Volkswagen Caddy teretni iz 2012. godine sa 1.6 TDI motorom i pređenih 301.000 km. Vozilo je korišćeno za dostavu.", "service_history": "Redovno servisiran na svakih 15.000 km; nedavno zamenjeni turbina i kvačilo.", "equipmentaj", "ABS", "servo volan", "radio", "pregradn [...]" |
| African QA | “system: You are a student sitting an oral examination in African literature. Answer the examiner's questions accurately and concisely. | user: Let us begin with Chinua Achebe. When was Things Fall Apart published, and in what setting? | ass [...]” | “The novel is built in three parts. The first two establish Igbo society on its own terms: its proverbs, festivals, legal procedures, religion and internal tensions, including Okonkwo's rigidity and the killing of Ikemefuna. Only in the third part does the missionary and colonial administration arrive. By giving roughly [...]” |
| Model: Gemini 3.1 Pro | ||
| General Customer Support | “How tall is the BaristaPro? I need to know if it fits under my cabinets.” | “The BaristaPro is 14.5 inches tall. Most standard kitchen cabinets have an 18-inch clearance, so it should fit perfectly.” |
{"cells":[{"text": "Serbian Car Ads","children":[]},{"text": ""Na prodaju VW Golf 5, 2007. godište, u odličnom stanju. Redovno održavan u ovlašćenom servisu. Od opreme poseduje automatsku klimu, tempomat i alu felne."", "children":[{"text":"VW Golf 5, 2007. godište, u odličnom stanju.","children":[]},{"text": "Redovno održavan u ovlašćenom servisu.","children":[]},{"text": "automatska klima","children":[]},{"text": "tempomat","children":[]},{"text": "alu felne","children":[]}]},{"text": "African QA","children":[]},{"text": ""user: Do we memorize publication dates for Achebe? | assistant: Just know the general era it was written. | user: So exact years aren't needed?"", "children":[{"text": "No, focus on the pre-colonial setting instead of publication years.","children":[]}]},{"text": "General Customer Support","children":[]},{"text": ""Does the ViewMax monitor include an HDMI cable?"", "children":[{"text": "Yes, the box includes one HDMI cable, a power adapter, the stand, and setup documentation. A DisplayPort cable is sold separately.","children":[]}]},{"text": "Hotel Reviews","children":[]},{"text": ""Write a balanced hotel review mentioning friendly staff, a convenient location, and aging amenities."", "children":[{"text": "The staff were genuinely friendly, and the location near shops and transit was extremely convenient. Our room was clean and comfortable, but the aging elevator, worn gym equipment, and dated bathroom need attention. Overall, it was a decent stay with room for improvement.","children":[]}]}]}
【譯註:此段落為表格內容,包含三個欄位:Dataset、Prompt 與 Completion。】
| Dataset | Prompt | Completion |
|---|---|---|
| Serbian Car Ads | “Volkswagen Golf 7 1.6 TDI, 2016. godište, 178.000 km. Uvezen iz Nemačke, bez ulaganja, redovno održavan. Comfortline paket, navigacija, tempomat i parking senzori.” | {"vehicle_description": "Golf 7 sa dizel motorom 1.6 TDI, 2016. godište i 178.000 km. Vozilo je uvezeno iz Nemačke, očuvano i spremno za vožnju bez dodatnih ulaganja.", "service_history": "Redovno servisiran; dostupna servisna dokumentacija iz Nemačke.", "equipment": ["Comfortline", "Navigacija", "Tempomat", "[...]]"} |
| African QA | “system: You are a student taking an exam on African history. | user: Teacher: What evidence shows that Great Zimbabwe was a major regional center? | assistant: Student: Its stone architecture and imported goods indicate wealth and trade. | [...]” | “Archaeologists found glass beads, Chinese ceramics, and coins linked to Indian Ocean commerce. Together with local gold-working evidence and monumental dry-stone walls, these finds show that Great Zimbabwe controlled resources and participated in long-distance trade rather than existing as an isolated settlement.” |
| Model: GLM-5.3 | ||
| Serbian Car Ads | “Toyota Yaris, 2012. godište, 1.0, 69 ks, 130.000 km. Gradski auto sa automatskom klimom, ekstremno ekonomičan. Servisiran u ovlašćenom servisu, knjižica kompletna.” | {"vehicle_description": "Toyota Yaris III, hečbek, 2012. godište, benzinski motor 1.0 VVT-i od 69 ks, prešao 130.000 km, bela boja, manuelni menjač sa 5 stepeni, prednji pogon, pouzdan gradski automobil", "service_history": "Kompletna servisna knjižica iz ovlašćenog servisa Toyota, ulje i filteri menjani na 126.000 km, [...]"} |
| African QA | “system: You are Professor Amara Diallo, a rigorous but encouraging teacher of African history conducting an oral exam. Give precise feedback on each answer, correct errors kindly, and end your turns with a follow-up question. | user: Good m [...]” | “Excellent. The pilgrimage also drew scholars, architects, and merchants toward Mali, boosting Timbuktu’s growth as a center of learning, and Mali appeared on European maps such as the Catalan Atlas of 1375. Here is a subtler question: why did Timbuktu flourish as a scholarly city, and what recent event in Mali threaten [...]” |
模型:Invent API
| Dataset | Prompt | Completion |
|---|---|---|
| General Customer Support | "Customer: 'My smart watch is not syncing heart rate data to the health app during workouts, even though the step counter works. I've unpaired and repaired the device three times. Does this sound like a sensor hardware failure or a software [...]" | "您好!我了解當健康資料無法同步時,特別是已經嘗試過多次配對與重新配對後,會有多麼令人沮喪。由於計步器功能正常,我們知道藍牙連線是活躍的,但我們需要判斷心率感測器本身是否故障,或是資料 [...]" |
| Hotel Reviews | "Write a two-sentence hotel assessment covering service, room quality, location, and amenities that balances both positive and negative points." | "禮賓服務無可挑剔且房間寬敞,然而,客房清潔時間不固定且 Wi-Fi 不穩定。雖然地理位置能提供無可比擬的市中心便利性,但缺乏功能完善的健身房以及吵雜的泳池區域,顯著降低了整體的住宿體驗。" |
{ "vehicle_description": "Renault Megane, godište 2015, pređeno 145.000 km. Vozilo je u dobrom stanju, redovno održavano.", "service_history": "Redovno servisiranje dokumentovano; veliki servis izvršen na 120.000 km.", "equipment-packages": ["Klima uređaj", "Navigacioni sistem", "Kožna sedišta"] } “The adoption of Christianity strengthened Aksum’s alignment with the Byzantine Empire, its primary Mediterranean trading partner, facilitating smoother commercial and diplomatic relations. However, this religious shift created a growing ideological rift with the Himyarite Kingdom in Southern Arabia. While their relatio [...]”
參考文獻
依原文順序編號,內文引用的數字對應下列條目。
- Adaption Labs. Invent a dataset. https://docs.adaptionlabs.ai/adaptive-data/invent-a-dataset/, 2026. Adaption API documentation. Accessed: 2026-09-24. Adaption Research Staff. AutoScientist: Automating the science of model training. Adaption Labs Blog, May 2026. URL https://adaptionlabs.ai/blog/autoscientist. Accessed: 2026-09-29. Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. A survey on data selection for language models. arXiv preprint arXiv:2402.16827, 2024. Anthropic. System card: Claude Opus 5. https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%20%20System%20Card.pdf, July 2026. Accessed: 2026-09-24. Quentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong, and Gauthier Gidel. On the stability of iterative retraining of generative models on their own data. In International Conference on Learning Representations (ICLR), 2024. arXiv:2310.00429. Meriem Boubdir, Edward Kim, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. Which prompts make the difference? data prioritization for efficient human llm evaluation. arXiv preprint arXiv:2310.14424, 2023a. Meriem Boubdir, Edward Kim, Beyza Ermis, Marzieh Fadaee, and Sara Hooker. Which prompts make the difference? data prioritization for efficient human llm evaluation, 2023b. URL https://arxiv.org/abs/2310.14424. Dan Brickley, Matthew Burgess, and Natasha Noy. Google dataset search: Building a search engine for datasets in an open web ecosystem.
- In The World Wide Web Conference (WWW), pp. 1365–1375, 2019. doi: 10.1145/3308558.3313685.
- Martin Briesch, Dominik Sobania, and Franz Rothlauf. Large language models suffer from their own output: An analysis of the self-consuming training loop. arXiv preprint arXiv:2311.16822, 2023. Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, and Paul Groth. Dataset search: A survey. The VLDB Journal, 29:251–272, 2020a. Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, and Paul Groth. Dataset search: a survey. The VLDB Journal, 29(1):251–272, 2020b. Everlyn Asiko Chimoto, Jay Gala, Orevaoghene Ahia, Julia Kreutzer, Bruce A Bassett, and Sara Hooker. Critical learning periods: Leveraging early training dynamics for efficient data pruning. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 9407–9426, 2024. DeepSeek-AI. DeepSeek-V4-Pro-0813. https://huggingface.co/deepseek-ai/DeepSeek-V4-PRO-0813, August 2026. Model card. Accessed: 2026-09-24. Elvis Dohmatob, Yunzhen Feng, and Julia Kempe. Model collapse demystified: The case of regression. In Advances in Neural Information Processing Systems (NeurIPS), 2024. arXiv:2402.07712. Google DeepMind. Gemini 3.1 Pro – model card. https://deepmind.google/models/model-car ds/gemini-3-1-pro/, February 2026. Accessed: 2026-09-24. Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, and Chloé Clavel. The curious decline of linguistic diversity: Training language models on synthetic text. In Findings of the Association for Computational Linguistics: NAACL 2024, pp. 3589–3604.
- Association for Computational Linguistics, 2024. URL https://aclanthology.org/2024.findings-naacl.228/. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751, 2019. Sara Hooker and Sudip Roy. Expand your world. Adaption Labs Blog, April 2026. URL https://adaptionlabs.ai/blog/expand-your-world. Accessed: 2026-09-30. Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, and Wei Wang. Followbench: A multi-level fine-grained constraints following benchmark for large language models, 2024. URL https://arxiv.org/abs/2310.20410. Emilia Kacprzak, Laura Koesten, Luis-Daniel Ibáñez, Tom Blount, Jeni Tennison, and Elena Simperl. Characterising dataset search queries. Companion of The Web Conference, 2019. Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, et al. The Stack: 3 TB of permissively licensed source code. arXiv preprint arXiv:2211.15533, 2022. Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton Earnshaw, Imran Haque, Sara M Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. Wilds: A benchmark of in-the-wild distribution shifts. In Marina Meila and Tong Zhang
- (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, pp. 5637–5664. PMLR, 18–24 Jul 2021. URL https://proceedings.mlr.press/v139/koh21a.html. Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, and Jason Weston. Autodata: An agentic data scientist to create high quality synthetic data, 2026. URL https://arxiv.org/abs/2606.25996. Keyu Li, Mohan Jiang, Dayuan Fu, Yunze Wu, Xiangkun Hu, Dequan Wang, and Pengfei Liu. Datasetresearch: Benchmarking agent systems for demand-driven dataset discovery. arXiv preprint arXiv:2508.06960, 2025a. Pengyue Li, Sheng Wang, Hua Dai, Zhiyu Chen, Zhifeng Bao, and Brian D. Davison. A survey on open dataset search in the llm era: Retrospectives and perspectives, 2025b. URL https://arxiv.org/abs/2509.00728. Rachel Lin, Bhavya Chopra, Wenjing Lin, Shreya Shankar, Madelon Hulsebos, and Aditya G Parameswaran. Rethinking dataset discovery with datascout. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, pp. 1–16, 2025. Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv preprint arXiv:2303.16634, 2023. Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. On LLMs-driven synthetic data generation, curation, and evaluation: A survey.
- arXiv preprint arXiv:2406.15126, 2024. Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, et al. A large-scale audit of dataset licensing and attribution in ai. Nature Machine Intelligence, 6(8):975–987, 2024. Marco Lui and Timothy Baldwin. langid.py: An off-the-shelf language identification tool. In Proceedings of the ACL 2012 system demonstrations, pp. 25–30, 2012. Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. WizardCoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568, 2023. Kelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze, and Sebastian Ruder. Understanding and mitigating language confusion in LLMs. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 6653–6677, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.380. URL https://aclanthology.org/2024.emnlp-main.380/. Max Marion, Ahmet Üstün, Luiza Pozzobon, Alex Wang, Marzieh Fadaee, and Sara Hooker. When less is more: Investigating data pruning for pretraining LLMs at scale. arXiv preprint arXiv:2309.04564, 2023.
- OpenAI. GPT-5.6 preview system card. https://deploymentsafety.openai.com/gpt-5-6-pre
- view, June 2026. Accessed: 2026-09-24. OpenRouter. OpenRouter: A unified interface for LLMs. https://openrouter.ai, 2026. Accessed:
- 2026-09-24. Vishakh Padmakumar and He He. Does writing with language models reduce content diversity?,
- URL https://arxiv.org/abs/2309.05196. Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar. Do ImageNet classifiers
- generalize to ImageNet? In Kamalika Chaudhuri and Ruslan Salakhutdinov (eds.), Proceedings
- of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine
- Learning Research, pp. 5389-5400. PMLR, 09-15 Jun 2019. URL https://proceedings.mlr. press/v97/recht19a.html. Yuji Roh, Geon Heo, and Steven Euijong Whang. A survey on data collection for machine learning:
- a big data-ai integration perspective. IEEE Transactions on Knowledge and Data Engineering,
- 33(4):1328-1347, 2019. Israfel Salazar, Manuel Fernández Burda, Shayekh Islam, Arshia Soltani Moakhar, Shivalika Singh,
- Fabian Farestam, Angelika Romanou, Danylo Boiko, Dipika Khullar, Mike Zhang, et al. Kalei-
- doscope: In-language exams for massively multilingual vision evaluation. In International Con-
- ference on Learning Representations, volume 2026, pp. 81112-81164, 2026. Max Schaffelder and Albert Gatt. Synthetic eggs in many baskets: The impact of synthetic data
- diversity on LLM fine-tuning. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David
- Jurgens (eds.), Findings of the Association for Computational Linguistics: ACL 2026, pp.
- 7265-
- 7293, San Diego, California, United States, July 2026. Association for Computational Linguistics. ISBN 978-8-89176-395-1. doi: 10.18653/v1/2026.findings-acl.360. URL https://aclanthology
- .org/2026.findings-acl.360/. Chantal Shaib, Venkata S. Govindarajan, Joe Barrow, Jiuding Sun, Alexa F. Siu, Byron C. Wallace,
- and Ani Nenkova. Standardizing the measurement of text diversity: A tool and a comparative
- analysis of scores, 2026. URL https://arxiv.org/abs/2403.00553. Luísa Shimabucoro, Sebastian Ruder, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. LLM
- see, LLM do: Leveraging active inheritance to target non-differentiable objectives. In Yaser
- Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on
- Empirical Methods in Natural Language Processing, pp. 9243-9267, Miami, Florida, USA, Novem-
- ber 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.521. URL https://aclanthology.org/2024.emnlp-main.521/. Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Ander-
- son. The curse of recursion: Training on generated data makes models forget. arXiv preprint
- arXiv:2305.17493, 2023. Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. Ai models collapse when trained on recursively generated data. Nature, 631(8022):755-759, 2024. doi: 10.1038/s41586-024-07566-y. Shivalika Singh, Freddie Vargus, Daniel D'souza, Börje Karlsson, Abinaya Mahendiran, Wei-Yin
- Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O'Mahony, Mike Zhang, Ramith
- Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergun, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Chien, Sebastian Ruder, Surya Guthikonda, Emad Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. Aya dataset: An open-access collection for multilingual instruction tuning. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11521–11567, Bangkok, Thailand, August 2024a. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.620. URL https://aclanthology.org/2024.acl-long.620/. Shivalika Singh, Freddie Vargus, Daniel D'souza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O'Mahony, et al. Aya dataset: An open-access collection for multilingual instruction tuning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 11521–11567, 2024b. Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al. Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 18761–18799, 2025.
- Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos. Beyond neural scaling laws: beating power law scaling via data pruning. Advances in Neural Information Processing Systems, 35:19523–19536, 2022. Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023. Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, 2022. Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton van den Hengel. On the value of out-of-distribution testing: An example of goodhart's law. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 407–417. Curran Associates, Inc., 2020. URL https://proceedings.neurips.cc/paper_files/paper/2020/file/045117b0e0a11a242b9765e79cbf113f-Paper.pdf. Megh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth, Sarath Chandar, and Partha Talukdar. Self-influence guided data reweighting for language model pre-training. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 2033–2045, 2023. Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos.
- D4: Improving llm pretraining via document de-duplication and diversification. Advances in Neural Information Processing Systems, 36:53983–53995, 2023.
- A. Torralba and A. A. Efros. Unbiased look at dataset bias. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition, CVPR '11, pp. 1521–1528, USA, 2011. IEEE Computer Society. ISBN 9781457703942. doi: 10.1109/CVPR.2011.5995347. URL https://doi.org/10.1109/CVPR.2011.5995347. Ke Wang et al. A survey on data synthesis and augmentation for large language models. arXiv preprint arXiv:2410.12896, 2024a. Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), pp. 13484–13508, 2023a. Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-Instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), 2023b. arXiv:2212.10560. Zifeng Wang, Chun-Liang Li, Vincent Perot, Long T. Le, Jin Miao, Zizhao Zhang, Chen-Yu Lee, and Tomas Pfister. CodecLM: Aligning language models with tailored synthetic data. In Findings of NAACL, 2024b. arXiv:2404.05875. Bosi Wen, Pei Ke, Xiaotao Gu, Lindong Wu, Hao Huang, Jinfeng Zhou, Wenchuang Li, Binxin Hu, Wendy Gao, Jiaxin Xu, Yiming Liu, Jie Tang, Hongning Wang, and Minlie Huang. Benchmarking complex instruction-following with multiple constraints composition, 2024. URL https://arxiv.org/abs/2407.03978. Kyle Wiggers.
- ‘open’ AI model licenses often carry concerning restrictions. https://techcrunch.com/2025/03/14/open-model-licenses-often-carry-concerning-restrictions/, March 2025. TechCrunch. Accessed: 2026-09-24. Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. Wizardlm: Empowering large pre-trained language models to follow complex instructions. In International Conference on Learning Representations, volume 2024, pp. 30745–30766, 2024. Jin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai, Huayang Li, and Jian Li. Learning to break the loop: Analyzing and mitigating repetitions for neural text generation. arXiv preprint arXiv:2206.02369, 2022. Z.ai. GLM-5.3. https://huggingface.co/zai-org/GLM-5.3, August 2026. Model card. Accessed: 2026-09-24. Biao Zhang, Zhongtao Liu, Colin Cherry, and Orhan Firat. When scaling meets llm finetuning: The effect of data, model and finetuning method, 2024. URL https://arxiv.org/abs/2402.17193. Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, et al. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176, 2025. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623, 2023.
- Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models, 2023. URL https://arxiv.org/abs/2311.07911.
- Yuchang Zhu, Huizhe Zhang, Bingzhe Wu, Jintang Li, Zibin Zheng, Peilin Zhao, Liang Chen, and Yatao Bian. Measuring diversity in synthetic datasets. arXiv preprint arXiv:2502.08512, 2025.
術語對照表
| 英文 | 中文 |
|---|---|
| Advantage | 優勢 |
| Agent | 代理人 |
| Base Model | 基礎模型 |
| Base model | 基底模型 |
| base model | 基礎模型 |
| Baseline | 基線 |
| Batch size | 批次大小 |
| Batch Size | 批次大小 |
| batch size | 批次大小 |
| Benchmark | 基準測試 |
| benchmark | 基準 |
| calibrate | 校準 |
| capabilities | 能力 |
| Collapse | 塌陷 |
| completion quality | 完成品質 |
| compressible | 可壓縮的 |
| Compression | 壓縮 |
| Compression Ratio | 壓縮比 |
| Configuration | 配置 |
| constrained | 受約束的 |
| constrained queries | 受約束的查詢 |
| constraints | 約束 |
| corpus | 語料庫 |
| Coverage | 覆蓋率 |
| curating data | 資料策展 |
| curation | 策展 |
| customer support | 客戶支援 |
| Data Curation | 資料策展 |
| data prioritization | 資料優先級排序 |
| data regime | 資料機制 |
| data variants | 資料變體 |
| dataset description | 資料集描述 |
| dataset retrieval agents | 資料集檢索代理人 |
| dataset scaling | 資料集規模化 |
| dataset search | 資料集搜尋 |
| de-duping | 去重複 |
| default values | 預設值 |
| demand-driven data curation | 需求驅動的資料策展 |
| distilling | 蒸餾 |
| Diversity | 多樣性 |
| diversity collapse | 多樣性塌陷 |
| Diversity Metrics | 多樣性指標 |
| diversity–quality trade-off | 多樣性與品質的權衡 |
| Domain | 領域 |
| Downstream | 下游 |
| end-point | 端點 |
| Environment | 環境 |
| exact-duplicate rate | 完全重複率 |
| Expert | 專家 |
| Filtering | 過濾 |
| fine | 罰金 |
| Fine-tuned | 微調的 |
| formatting constraints | 格式約束 |
| frontier LLMs | 前沿 LLM |
| Frontier Model | 前沿模型 |
| Frontier model | 前沿模型 |
| GEMM | GEMM |
| Generation | 生成 |
| generation strategies | 生成策略 |
| generators | 生成器 |
| Held-out | 保留 |
| human-annotated | 人工標記的 |
| Hyperparameter | 超參數 |
| Hyperparameters | 超參數 |
| imitation | 模仿 |
| Judge | 裁判 |
| language confusion | 語言混淆 |
| language drift | 語言偏移 |
| Language Model | 語言模型 |
| Large Language Models (LLMs) | 大型語言模型 |
| Level | 關卡 |
| lexical diversity | 詞彙多樣性 |
| lexical prompt diversity | 詞彙提示多樣性 |
| LLM Judge | LLM 評審 |
| LM Judge | 語言模型評審 |
| maximization-based decoding | 基於最大化的解碼 |
| medical QA | 醫療問答 |
| Metadata | 中介資料 |
| model families | 模型家族 |
| multilingual setting | 多語言設定 |
| multiple choice output restrictions | 多選輸出限制 |
| N-gram | N 元語法 |
| n-gram diversity | n-gram 多樣性 |
| on-specification | 符合規格的 |
| on-the-fly | 即時 |
| Open-weight | 開放權重 |
| open-weight model | 開放權重模型 |
| Open-weights Model | 開放權重模型 |
| output token limit | 輸出 Token 限制 |
| output tone | 輸出語調 |
| pareto-frontier | 帕累托前沿 |
| Performance | 效能 |
| Persona | 人設 |
| Pipeline | 管線 |
| Plateau | 收斂高原 |
| Post-training | 後訓練 |
| Post-Training | 後訓練 |
| Prompt | 提示 |
| Prompting | 提示 |
| Proprietary | 專有的 |
| proprietary | 專有 |
| Proprietary Model | 專有模型 |
| Pruning | 剪枝 |
| quality gains | 品質提升 |
| Query | 查詢 |
| Rate | 速率 |
| Reasoning | 推理 |
| redundant consecutive repetitions | 冗餘的連續重複 |
| relative gains | 相對增益 |
| Relevance | 相關性 |
| repetition | 重複 |
| repetitive | 重複的 |
| Representation | 代表性 |
| Return | 報酬 |
| Robustness | 穩健性 |
| Router | 路由器 |
| samples | 樣本 |
| Seed | 種子 |
| seed pool | 種子池 |
| semantic content | 語義內容 |
| Specification | 規格 |
| state | 狀態 |
| Stem | 主幹 |
| Synthetic | 合成的 |
| Synthetic Data | 合成資料 |
| Synthetic Data Generation | 合成資料生成 |
| target language | 目標語言 |
| task types | 任務類型 |
| task-specific metrics | 任務特定指標 |
| temperature | 溫度 |
| Temperature | 溫度 |
| terms of service | 服務條款 |
| Token | Token |
| tone constraints | 語調約束 |
| Top-1 rank share | Top-1 排名佔比 |
| trade-off | 權衡 |
| training | 訓練 |
| Training data | 訓練資料 |
| training objective | 訓練目標 |
| Turn | 回合 |
| un-tuned | 未微調的 |
| unconstrained | 不受約束的 |
| unconstrained queries | 無約束查詢 |
| unconstrained setting | 無約束設定 |
| upper bound | 上限 |
| World Model | 世界模型 |
| World model | 世界模型 |
| zero data regime | 零資料機制 |
| zero training data regime | 零訓練資料機制 |