Cosmos:面向物理 AI 的世界基礎模型平台
原始論文:Cosmos World Foundation Model Platform for Physical AI 作者:NVIDIA arXiv ID:2501.03575v3 日期:2025-01-07(v1 初版;v3 修訂於 2025-07-09) 標籤:World Model Physical AI 影片生成 Tokenizer Diffusion Autoregressive 機器人
目錄
摘要
物理 AI (Physical AI) 需要先以數位方式進行訓練。它需要一個自身的數位孿生 (Digital Twin)——即策略模型 (Policy Model),以及一個世界的數位孿生——即世界模型 (World Model)。本文提出 Cosmos 世界基礎模型 (World Foundation Model, WFM) 平台,協助開發者為其物理 AI 系統打造客製化的世界模型。我們將世界基礎模型定位為一種通用的世界模型,可以微調 (Fine-tune) 為下游應用的客製化世界模型。我們的平台涵蓋影片整理流程 (Video Curation Pipeline)、預訓練的世界基礎模型、預訓練世界基礎模型的後訓練 (Post-training) 範例,以及影片 tokenizer。為了協助物理 AI 開發者解決社會上最關鍵的問題,我們將 Cosmos 開源,模型權重亦以寬鬆授權開放,可經由 NVIDIA Cosmos-Predict1 取得。
1 引言
物理 AI 是一種配備感測器 (Sensor) 與致動器 (Actuator) 的 AI 系統:感測器讓它能觀察世界,致動器讓它能與世界互動並改變世界。它有望將人類從危險、繁重或乏味的體力工作中解放出來。過去十年,多個 AI 領域因資料與運算規模化而大幅進展,物理 AI 卻僅緩步前行。這主要是因為擴充物理 AI 的訓練資料困難得多——所需的資料必須包含交錯的觀察與動作序列。這些動作會擾動物理世界,可能對系統與世界造成嚴重損害;當 AI 尚處於必須大量探索的初期階段時尤其如此。世界基礎模型——一個可讓物理 AI 安全互動的物理世界數位孿生——長久以來被視為解決資料規模化問題的解方。
圖 1(原文為影片,以下為各影片的代表影格):








圖 1:Cosmos 世界基礎模型。 預訓練的 Cosmos WFM 能生成高品質、3D 一致且物理正確的影片。Cosmos 模型套件同時包含擴散 (Diffusion) 與自迴歸 (Autoregressive) transformer 模型,分別以影片的連續與離散潛在表徵 (Latent Representation) 進行訓練。以專門資料集對這些 WFM 進行後訓練後,即可應用於各式各樣的物理 AI 場景。具體而言,我們展示了具相機可控性的模型、可依指令執行機器人操作的模型,以及自動駕駛情境的模型。完整影片與更多範例請造訪我們的網站。

圖 2: 預訓練 WFM 是世界模型的通才 (Generalist),以大規模、多樣化、涵蓋真實世界物理各個面向的影片資料集訓練而成。這些預訓練世界基礎模型可透過後訓練特化到目標物理 AI 場景。後訓練資料集通常是從目標物理 AI 場景蒐集的「提示詞 (Prompt)」-影片配對,提示詞可以是動作指令、軌跡、指示等形式。由於預訓練 WFM 提供了良好的基礎,後訓練所需的資料集可以小得多。這種「先預訓練、再後訓練」的模式是打造物理 AI 系統的高效策略。圖中虛線代表資料迴圈。
本文介紹用於打造物理 AI 的 Cosmos 世界基礎模型平台。我們主要關注視覺世界基礎模型,其中觀察以影片形式呈現,而擾動 (Perturbation) 可以以多種形式存在。如圖 2 所示,我們提出「先預訓練、再後訓練」的範式,將 WFM 分為預訓練 WFM 與後訓練 WFM。為了建立預訓練 WFM,我們利用大規模影片訓練資料集,讓模型接觸多樣的視覺經驗,成為通才。為了建立後訓練 WFM,我們使用從特定物理 AI 環境蒐集的資料集微調預訓練 WFM,得到針對目標場景特化的 WFM。圖 1 展示了我們的預訓練與後訓練 WFM 的範例結果。
資料決定了 AI 模型的上限。為了打造上限更高的預訓練 WFM,我們開發了影片資料整理流程 (Data Curation Pipeline),用它找出影片中動態豐富、視覺品質高、有助於學習視覺內容所編碼物理知識的片段。我們透過此流程,從長達 2,000 萬小時的影片庫中萃取出約 1 億段長度介於 2 到 60 秒的影片片段。對每個片段,我們使用視覺語言模型 (Visual Language Model, VLM) 每 256 影格產生一段影片描述。影片處理是計算密集型工作,我們利用現代 GPU 內建的 H.264 影片編解碼器硬體實作進行解碼與轉檔。我們的影片資料整理流程使用許多預訓練的圖像/影片理解模型,這些模型的吞吐量各異。為了讓可訓練影片資料的整體產出吞吐量最大化,我們建構了基於 Ray 的協調流程 (Moritz et al. 2017)。細節見第 3 節。
我們探索兩種可規模化的預訓練 WFM 建構方法,詳見第 5 節:基於 transformer 的擴散模型 (Diffusion Model) 與基於 transformer 的自迴歸模型 (Autoregressive Model)。擴散模型透過從高斯雜訊影片中逐步去除雜訊來生成影片;自迴歸模型則依照預設順序、以過去的生成為條件,逐片段生成影片。兩種方法都將困難的影片生成問題分解為較容易的子問題,使其更易處理。我們採用最先進的 transformer 架構以取得可擴展性。第 5.1 節提出一種基於 transformer 的擴散模型設計,展現強大的世界生成能力;第 5.2 節提出一種基於 transformer 的自迴歸模型設計用於世界生成。
基於 transformer 的擴散模型與自迴歸模型都以 token 作為影片的表徵,前者使用向量形式的連續 token,後者使用整數形式的離散 token。我們注意到,影片的 token 化 (Tokenization)——將影片轉換為一組 token 的過程——絕非易事。影片蘊含關於視覺世界的豐富資訊,然而為了讓 WFM 易於學習,我們必須把影片壓縮成精簡的 token 序列,同時盡可能保留原始內容,因為世界基礎模型訓練的計算複雜度會隨 token 數量增長。就許多層面而言,打造影片 tokenizer 與打造影片編解碼器 (Video Codec) 相似。我們開發了基於注意力機制 (Attention Mechanism) 的編碼器-解碼器架構,用於學習連續與離散兩種 token 的影片 token 化,詳見第 4 節。
在第 6 節,我們微調預訓練 WFM,得到面向各種物理 AI 任務的後訓練 WFM。第 6.1 節將預訓練擴散 WFM 微調為以相機姿態 (Camera Pose) 為條件的模型。這種後訓練創造了一個可導覽的虛擬世界,使用者可以移動虛擬視角來探索生成的世界。第 6.2 節在多種由影片-動作序列構成的機器人任務上微調我們的 WFM。我們證明利用預訓練 WFM,能更準確地依據機器人採取的動作預測世界的未來狀態。第 6.3 節展示如何為各種自動駕駛相關任務微調預訓練 WFM。
我們開發的 WFM 預期用途是服務物理 AI 開發者。為了讓開發者更安全地使用世界基礎模型,我們開發了強大的護欄 (Guardrail) 系統,包含阻擋有害輸入的 pre-Guard 與阻擋有害輸出的 post-Guard,細節見第 7 節。
我們的目標是打造一個世界基礎模型平台,協助物理 AI 開發者推進他們的系統。為此,我們將預訓練世界基礎模型與 tokenizer 以 NVIDIA 開放模型授權 (NVIDIA Open Model License) 於 NVIDIA Cosmos 開放。儘管本文在世界基礎模型設計上做出多項改進,世界基礎模型問題仍遠未解決,需要更多研究來進一步推進最先進水準。
2 世界基礎模型平台
令 $x_{0:t}$ 為從時間 $0$ 到 $t$ 對真實世界的一系列視覺觀察,令 $c_{t}$ 為對世界的擾動。如圖 3 所示,WFM 是一個模型 $\mathcal{W}$,它根據過去的觀察 $x_{0:t}$ 與當前的擾動 $c_{t}$,預測時間 $t+1$ 的未來觀察 $\hat{x}_{t+1}$。在我們的設定中,$x_{0:t}$ 是 RGB 影片,而 $c_{t}$ 是可採取多種形式的擾動:可以是物理 AI 採取的動作、隨機擾動、對擾動的文字描述等。
圖 3: 世界基礎模型 $\mathcal{W}$ 是根據過去觀察 $x_{0:t}$ 與當前擾動 $c_{t}$ 生成世界未來狀態 $x_{t+1}$ 的模型。【譯註:原圖為示意圖,HTML 版未提供對應圖檔。】
2.1 未來的 Cosmos
我們相信 WFM 對物理 AI 開發者有多方面的用處,包括(但不限於):
- 策略評估 (Policy Evaluation)。 指評估物理 AI 系統中策略模型的品質。與其將訓練好的策略部署到真實世界運行的物理 AI 系統上進行評估,不如讓物理 AI 系統的數位副本與世界基礎模型互動。基於 WFM 的評估更省成本也更省時間。藉助 WFM,開發者可以把策略模型部署到原本無法取得的未見環境。WFM 能協助開發者快速排除不合格的策略,把實體資源集中在少數有希望的策略上。
- 策略初始化 (Policy Initialization)。 策略模型根據當前觀察與給定任務,生成物理 AI 系統要採取的動作。訓練良好的 WFM 依據輸入擾動對世界的動態模式建模,可作為策略模型的良好初始化,有助於解決物理 AI 的資料稀缺問題。
- 策略訓練 (Policy Training)。 WFM 搭配獎勵模型 (Reward Model),可在強化學習 (Reinforcement Learning) 設定中作為物理世界的代理,向策略模型提供回饋。代理人 (Agent) 可以透過與 WFM 互動,逐漸熟練地解決任務。
- 規劃或模型預測控制 (Model-Predictive Control)。 WFM 可用來模擬物理 AI 系統採取不同動作序列後的不同未來狀態,再由成本/獎勵模組根據結果量化這些動作序列的表現。物理 AI 便能依據模擬結果整體執行最佳動作序列(如規劃演算法),或以滾動時域 (Receding Horizon) 方式執行(如模型預測控制)。世界模型的準確度是這些決策策略表現的上限。
- 合成資料生成 (Synthetic Data Generation)。 WFM 可用於生成訓練用的合成資料,也可以微調為以深度圖或語意圖等渲染中繼資料為條件的模型。條件式 WFM 可用於 Sim2Real 應用場景。
雖然我們列出這些可能性,本文並未包含將 Cosmos WFM 應用於上述用途的實證結果。我們期待在未來的工作中驗證這些主張。
2.2 當前的 Cosmos

圖 4:Cosmos 世界基礎模型平台由幾個主要元件組成:影片整理器 (Video Curator)、影片 tokenizer、預訓練世界基礎模型、世界基礎模型後訓練範例,以及護欄。
圖 4 呈現本文中 Cosmos WFM 平台所提供的內容,包括影片整理器、影片 token 化、世界基礎模型預訓練、世界基礎模型後訓練,以及護欄。
影片整理器。 我們開發了可規模化的影片資料整理流程。每部影片會被切分成不含場景轉換的個別鏡頭 (Shot),接著對這些片段套用一系列過濾步驟,找出高品質且動態資訊豐富的子集供訓練使用。這些高品質鏡頭再由 VLM 進行標註,最後執行語意去重 (Semantic Deduplication),建構出多樣而精簡的資料集。
影片 token 化。 我們開發了一系列具有不同壓縮率的影片 tokenizer。這些 tokenizer 具因果性 (Causal)——當前影格的 token 計算不依賴未來的觀察。這種因果設計有多項好處。在訓練面,它讓圖像與影片的聯合訓練成為可能,因為當輸入是單張圖像時,因果影片 tokenizer 同時也是圖像 tokenizer。這一點很重要,能讓影片模型利用圖像資料集訓練——圖像資料蘊含豐富的世界外觀資訊,且通常更多樣。在應用面,因果影片 tokenizer 與生存在因果世界中的物理 AI 系統更加契合。
WFM 預訓練。 我們探索兩種可規模化的預訓練世界基礎模型建構方法——擴散模型與自迴歸模型,並採用 transformer 架構以取得可擴展性。
對基於擴散的 WFM,預訓練包含兩個步驟:1) Text2World 生成預訓練,2) Video2World 生成預訓練。具體來說,我們先訓練模型根據輸入文字提示詞生成影片世界,再微調模型使其根據過去影片與輸入文字提示詞生成未來的影片世界——我們稱之為 Video2World 生成任務。
對基於自迴歸的 WFM,預訓練包含兩個步驟:1) 基本的下一個 token 生成,2) 文字條件式 Video2World 生成。我們先訓練模型根據過去影片的輸入生成未來的影片世界——即前瞻生成 (Foresight Generation),再微調模型使其根據過去影片與文字提示詞生成未來的影片世界。
Video2World 生成模型是一種預訓練世界模型,根據當前觀察(過去影片)與控制輸入(提示詞)生成未來。對基於擴散與基於自迴歸的 WFM,我們都建立了不同容量的模型家族,並研究其在各種下游應用的效果。
我們進一步微調預訓練擴散 WFM,得到一個擴散解碼器 (Diffusion Decoder),以強化自迴歸模型的生成結果。為了更好地控制 WFM,我們也基於大型語言模型 (Large Language Model, LLM) 建立了提示詞上採樣器 (Prompt Upsampler)。
世界模型後訓練。 我們展示預訓練 WFM 在多個下游物理 AI 應用的成果。我們以相機姿態作為輸入提示詞微調預訓練 WFM,讓使用者能在生成的世界中自由導覽。我們也展示了預訓練 WFM 如何針對人形機器人 (Humanoid) 與自動駕駛任務進行微調。
護欄。 為了安全使用所開發的世界基礎模型,我們開發了阻擋有害輸入與輸出的護欄系統。
3 資料整理
本節描述我們的影片整理流程 (Video Curation Pipeline),它為 tokenizer 與 WFM 產出高品質的訓練資料集。如圖 5 所示,流程包含五個主要步驟:1) 切分 (Splitting)、2) 過濾 (Filtering)、3) 標註 (Annotation)、4) 去重 (Deduplication)、5) 分片 (Sharding)。每個步驟都經過調校,以提升資料品質並滿足模型訓練的需求。我們先介紹原始資料集,再逐一詳述各步驟。

圖 5:Cosmos 影片整理器包含五大步驟:1) 切分、2) 過濾、3) 標註、4) 去重、5) 分片。切分步驟把長影片分割為鏡頭並轉檔為片段;過濾步驟移除對世界基礎模型建構價值不高的片段;標註步驟為每個片段加上影片描述,之後片段存入影片片段資料庫。要取得訓練資料集,先執行語意去重,再依解析度與長寬比對影片片段進行分片。
3.1 資料集
我們同時使用專有影片資料集與公開的開放領域網路影片來訓練模型。我們的目標是賦能物理 AI 開發者。為此,我們整理的影片訓練資料集涵蓋各種物理 AI 應用,並鎖定下列影片類別:
- 駕駛(11%)
- 手部動作與物件操作(16%)
- 人體動作與活動(10%)
- 空間感知與導航(16%)
- 第一人稱視角(8%)
- 自然動態(20%)
- 動態相機運動(8%)
- 合成渲染(4%)
- 其他(7%)
這些影片廣泛涵蓋不同的視覺物件與動作,其多樣性可提升 WFM 的泛化能力,並幫助模型應對不同的下游任務。這些影片的非結構化特性與龐大數量,使得高效處理在演算法與基礎設施層面都充滿挑戰。影片可能以各式各樣的編解碼器編碼,長寬比、解析度、長度等各不相同。許多影片還經過後製或加入各種視覺特效,若處理不當,可能在生成影片中引入不必要的瑕疵 (Artifact),並損害世界模型的表現。
我們總共累積了約 2,000 萬小時的原始影片,解析度從 720p 到 4K。然而,相當大比例的影片資料要嘛語意上重複,要嘛不含學習世界物理所需的有用資訊。因此,我們設計了一系列資料處理步驟,從原始影片中找出最有訓練價值的部分。我們也蒐集了圖像資料,因為圖像與影片的聯合訓練已被證明能提升生成影片的視覺品質並加速模型訓練。得益於資料整理流程的模組化設計,我們可以用它同時處理圖像與影片資料,為預訓練與微調產生資料集。我們為預訓練產生約 $10^{8}$ 個影片片段,為微調產生約 $10^{7}$ 個。
3.2 切分
我們的影片長度不一,而現代深度學習模型無法直接處理非常長的影片。此外,許多影片包含鏡頭轉換 (Shot Transition):影片可能從一個場景切換到完全無關的另一個場景,例如從紐約市現代廚房中兩人交談的畫面,切換到非洲草原上獅子追逐斑馬的場景。依據鏡頭變化分割影片、產生視覺上連貫的影片片段非常重要,這樣模型才能學習物理上合理的視覺內容轉變,而非人工剪輯出來的轉場。
3.2.1 鏡頭偵測
切分的目標是把任意長度的原始影片在時間軸上分割為不含鏡頭變化的片段。它以原始影片為輸入,輸出每個鏡頭的起始與結束影格索引。短於 2 秒的片段會被捨棄,因為它們可能是轉場或視覺特效;長於 60 秒的片段會再切分,使最大長度為 60 秒。後續的過濾步驟則判斷片段是否包含學習世界物理的有用資訊。
鏡頭邊界偵測 (Shot Boundary Detection) 是經典的電腦視覺問題。既有方法根據視覺特徵空間的變化偵測鏡頭邊界,差別在於如何從影格學習視覺特徵。我們在表 1 中評估了幾種演算法:PySceneDetect (Castellano 2024)、Panda70M (Chen et al. 2024)、TransNetV2 (Soucek and Lokoc 2024) 與 AutoShot (Zhu et al. 2023)。
PySceneDetect 是熱門函式庫,透過對 HSV 色彩空間中色彩直方圖的時間變化設定閾值來偵測鏡頭變化,近期的 MovieGen (Polyak et al. 2024) 也採用了它。Panda70M 在 PySceneDetect 之上加入基於 CLIP 嵌入 (Embedding) 的拼接與過濾。TransNetV2 與 AutoShot 則是神經網路方法,在 100 影格的滑動輸入視窗下預測每一影格是轉場影格的機率。
選擇能妥善處理重度剪輯影片的演算法至關重要,因為這類影片常有複雜的鏡頭變化並疊加各種視覺特效。這促使我們建立專門的基準測試,評估各方法能否從影片中產生鏡頭切割乾淨的片段。我們的基準(命名為 ShotBench 【腳註:ShotBench 見 https://github.com/NVlabs/ShotBench。】)包含既有資料集,如 RAI、BBC Planet Earth (AI Image Lab, University of Modena 2016)、ClipShots (Tang et al. 2018) 與 SHOT (Zhu et al. 2023)。對 ClipShots,我們把轉場影格定義為每段鏡頭標註起訖點的中點,以與其他資料集保持一致。
| 資料集 | 指標 | PySceneDetect | Panda70M | TransNetV2 | AutoShot |
|---|---|---|---|---|---|
| BBC | Precision $\uparrow$ | 0.894 | 0.959 | 0.983 | 0.984 |
| Recall $\uparrow$ | 0.884 | 0.653 | 0.951 | 0.922 | |
| F1 $\uparrow$ | 0.889 | 0.777 | 0.967 | 0.952 | |
| RAI | Precision $\uparrow$ | 0.856 | 0.933 | 0.918 | 0.889 |
| Recall $\uparrow$ | 0.807 | 0.746 | 0.921 | 0.923 | |
| F1 $\uparrow$ | 0.831 | 0.829 | 0.919 | 0.906 | |
| SHOT | Precision $\uparrow$ | 0.769 | 0.949 | 0.883 | 0.866 |
| Recall $\uparrow$ | 0.673 | 0.462 | 0.767 | 0.804 | |
| F1 $\uparrow$ | 0.718 | 0.622 | 0.821 | 0.834 | |
| ClipShots | Precision $\uparrow$ | 0.395 | 0.649 | 0.685 | 0.653 |
| Recall $\uparrow$ | 0.602 | 0.424 | 0.772 | 0.781 | |
| F1 $\uparrow$ | 0.477 | 0.513 | 0.726 | 0.711 |
表 1: 各切分演算法在不同資料集上的比較。
表 1 在 ShotBench 上比較各方法。TransNetV2 與 AutoShot 的信心閾值皆設為 0.4。對 Panda70M,為求公平比較,我們遵循其切分實作但不含過濾步驟。端到端學習式方法(如 TransNetV2 與 AutoShot)表現遠優於使用手工特徵或啟發式規則的方法(如 PySceneDetect 與 Panda70M)。儘管 TransNetV2 與 AutoShot 在既有資料集上表現相當,我們發現 TransNetV2 在更具挑戰性的鏡頭變化上表現更佳。使用端到端神經網路(即 TransNetV2)也讓我們能利用現代 GPU 加速來提升切分吞吐量,避免像 Panda70M 這類混合方法以複雜邏輯結合 PySceneDetect 與 ImageBind 嵌入 (Girdhar et al. 2023) 所帶來的麻煩。
3.2.2 轉檔
我們的影片使用許多不同的編解碼器與各種設定,為資料整理帶來挑戰。我們把鏡頭偵測產出的每個影片片段重新編碼為一致的高品質 mp4 格式,簡化後續的資料整理流程。統一影片編解碼器後,模型訓練用資料載入器 (Dataloader) 的穩定性與效率也大幅提升。我們使用高位元率的 h264_nvenc 編解碼器,並以快速運動與高頻紋理的影片對設定做壓力測試,確保沒有可察覺的視覺劣化。
我們在表 2 中全面評估不同的硬體與軟體轉檔配置,以最大化吞吐量。現代 GPU 提供硬體加速的影片編解碼能力:NVIDIA L40S 同時具備解碼 (NVDEC) 與編碼 (NVENC) 硬體加速器,而 NVIDIA H100 只有 NVDEC。為了與 L40S 公平比較,我們在表 2 中讓 H100 使用最多可用的 CPU 核心(28 個而非 1 個)作為補償。L40S 的吞吐量比 H100 高約 17%(0.0674 vs 0.0574)。在軟體配置上,從 libx264 換成 h264_nvenc,以及對同一部影片的多個片段做批次轉檔,都能顯著提升吞吐量。我們觀察到 ffmpeg 難以充分利用 NVDEC/NVENC 加速器,在多 GPU 節點上尤其明顯。把影片串流轉檔從 ffmpeg 換成 PyNvideoCodec 後,加速器使用率大幅提升,帶來最大的吞吐量改善(0.3702 vs 0.1026)。我們僅保留 ffmpeg 做音訊重混 (Audio Remixing),並使用 PyNvideoCodec 以更好地發揮 GPU 的運算能力。結合所有改進後,吞吐量提升約 $6.5\times$。
| 方法 | GPU | CPU | 編解碼器 | 批次 | NVDEC(加速器數) | NVENC(加速器數) | 吞吐量(影片/秒) |
|---|---|---|---|---|---|---|---|
| ffmpeg | H100 | 28 | libx264 | 1 | 7 | 0 | 0.0574 |
| ffmpeg | L40S | 1 | h264_nvenc | 1 | 3 | 3 | 0.0674 |
| ffmpeg | L40S | 1 | h264_nvenc | 16 | 3 | 3 | 0.1026 |
| pynvc+ffmpeg | L40S | 1 | h264_nvenc | 1 | 3 | 3 | 0.3702 |
表 2: 不同軟體設定下的轉檔效能。
3.3 過濾
切分步驟產出的影片片段品質參差、主題各異、雜訊不少。我們設計過濾步驟以:1) 移除視覺品質達不到最低要求的影片片段,2) 挑選適合微調的高品質影片片段,3) 為建構 WFM 調整資料分布。我們透過運動過濾、視覺品質過濾、文字過濾與影片類型過濾來達成上述目標。
3.3.1 運動過濾
運動過濾有兩個主要目標:1) 移除靜態或帶有隨機突兀相機運動(通常來自手持相機)的影片;2) 為影片標記不同類型的相機運動(如平移 (Pan)、變焦 (Zoom)、俯仰 (Tilt) 等),為模型訓練提供額外指引資訊。
我們建立了一個輕量級分類器做運動過濾。分類器的輸入是從影片片段萃取的運動向量 (Motion Vector) 或光流 (Optical Flow) 序列。分類器基於 ViT 架構,以有標籤的影片訓練。我們試驗了 h264 編解碼器的運動向量、Farneback 光流演算法 (Farnebäck 2003),以及 NVIDIA TensorRT 加速的光流估計網路。我們發現以 NVIDIA TensorRT 加速光流估計為基礎的分類器效果最好,在運動過濾上有很高的分類準確率。
3.3.2 視覺品質過濾
我們以失真 (Distortion) 與外觀品質兩項標準做視覺品質過濾。首先,我們移除帶有失真的影片片段,如瑕疵、雜訊、模糊、低銳利度、過曝、欠曝等。我們使用基於 DOVER (Wu et al. 2023a)、以人工評分影片訓練的影片品質評估模型,為每個片段給出感知品質分數,並據此移除分數位於最後 $15\%$ 的片段。其次,我們過濾外觀品質低的影片片段:對輸入片段取樣影格,套用圖像美學模型 (Schuhmann 2022)。由於美學對物理 AI 而言相對次要,我們設定了保守的閾值,即 $3.5$。
3.3.3 文字覆疊過濾
部分影片經過後製加上文字,向觀眾提供額外資訊。我們也發現文字往往與各種視覺特效同時出現。我們的目標是學習世界的物理,因此移除帶有過多此類文字的影片至關重要。注意我們針對的是後製加入的文字,而非影片原始場景中的文字,例如駕駛影片中的路牌街名。
我們訓練了一個基於 MLP 的二元分類器來偵測這類影片。分類器的輸入是以 InternVideo2 (Wang et al. 2025) 萃取的影片嵌入。我們使用專有 VLM 建立訓練集,標記正、負樣本影片。訓練出的模型在驗證集上達到很高的預測準確率。
3.3.4 影片類型過濾
為了調整訓練資料分布並過濾不想要的影片類型,我們設計了一套完整的分類法 (Taxonomy),依內容類型與視覺風格對影片分類,並訓練分類器為每個影片片段標上分類法中的類別。我們排除可能導致生成品質差或動態不真實的特定影片類型(如抽象視覺圖案、電玩畫面、動畫內容等)以精煉資料。我們進一步調整資料分布:對與 WFM 更相關的類別(如人類動作、人與物件互動等)上採樣 (Upsampling),對較不重要的類別(如自然或風景影片)下採樣 (Downsampling)。
由於沒有符合我們分類法的現成標註資料集,我們利用專有 VLM 為分類器建立訓練與評估資料。對每個影片片段,我們以八個均勻取樣的影格提示 VLM,詢問最合適的分類標籤。利用標註資料,我們在與文字過濾相同的 InternVideo2 嵌入上訓練 MLP 分類器。
3.4 標註
文字描述通常與圖像及影片資料配對,為世界模型訓練提供監督訊號與條件。我們使用 VLM 為每個影片片段生成高品質且一致的描述 (Caption)。我們對 VLM 做了設定,使其專注於影片中的客觀事實與細節。用這種方式提供影片描述、而非依賴替代文字 (Alt Text),也減輕了世界模型的學習負擔,因為訓練時不必適應不同的文字風格或格式。
我們在自家影片上測試了多個最先進 (SOTA) 的描述生成方法(即 VFC (Ge et al. 2024)、Qwen2-VL (Wang et al. 2024b)、VILA (Lin et al. 2024b; Xue et al. 2024)),根據小規模人工評估,VILA 生成的描述更準確。我們使用內部的 13B 參數 VILA 模型,並針對影片描述任務微調。它具備適合處理長多影格上下文的加大上下文視窗,最大輸入與輸出 token 長度分別為 5904 與 256。為了提升推論效率,我們使用 FP8 量化的 TensorRT-LLM 引擎,相較於 PyTorch 半精度基線,吞吐量提升 10 $\times$,如表 3 所示。我們以「Elaborate on the visual and narrative elements of the video in detail」提示 VILA,並輸入從片段均勻取樣的 8 個影格。描述的平均長度為 559 個字元或 97 個單詞。
| 引擎 | 精度 | 批次大小 | 吞吐量(片段/秒) | 吞吐量(token/秒) |
|---|---|---|---|---|
| PyTorch | FP16 | 1 | 0.21 | 49.6 |
| TRT-LLM | FP16 | 1 | 0.40 | 95.6 |
| TRT-LLM | FP16 | 16 | 1.09 | 260.9 |
| TRT-LLM | FP8 | 16 | 1.96 | 470.6 |
表 3: VILA 在單張 H100 GPU 上的推論吞吐量比較。
3.5 去重
鑑於影片數量龐大,訓練集中可能存在重複或近似重複的樣本。去重對建立更均衡、多樣的資料分布至關重要,同時也能提升訓練效率、降低記憶特定訓練樣本的風險。
我們採用 SemDeDup (Abbas et al. 2023) 與 DataComp (Gadre et al. 2024) 的做法進行可規模化的語意去重。我們重用過濾階段計算的 InternVideo2 嵌入,並以多節點 GPU 加速的 k-means 實作 (RAPIDS 2023)($k=10{,}000$)對嵌入分群。我們在每個嵌入群內計算成對距離以找出重複項。偵測到重複影片時,我們保留解析度最高的那一部,確保去重不損失品質。為了避免把整個成對距離矩陣存放在 GPU 記憶體中,我們以 256 為區塊、即時計算所需的上三角矩陣與 argmax 歸約。去重階段移除了約 $30\%$ 的訓練資料。
我們也利用萃取的 InternVideo2 嵌入與分群結果,建立了一個支援以自由文字與影片查詢整個訓練資料集的視覺搜尋引擎。這個搜尋引擎對除錯資料問題、理解預訓練資料集與下游應用之間的差距很有幫助。
3.6 分片
此步驟的目標是把處理好的影片片段打包成模型訓練器可直接使用的 webdataset。我們依解析度、長寬比與長度對影片分片,以配合我們的訓練課程 (Training Curriculum)。除了預訓練資料集,我們也利用上述各種過濾器建立品質更高的微調資料集。
3.7 基礎設施
我們的資料處理基礎設施使用 AnyScale Ray (Moritz et al. 2017) 實作串流管線系統,服務地理上分散的叢集,解決大規模 ML 工作流程的兩大挑戰:同質節點間的高效資源利用,以及對高延遲資料來源連線的穩健運作。透過把資料傳輸與運算解耦,管線能高效地搭配遠端資料儲存運作,且記憶體需求隨管線複雜度而非資料集大小成長,實現無上限的串流處理。
我們的架構透過平行管線階段同時利用互補的硬體資源,例如同時使用網路頻寬做資料擷取、NVDEC 單元做影片解碼、GPU 做計算密集的轉換。我們擴充了 Fragmentation Gradient Descent 演算法 (Weng et al. 2023) 來最佳化這種多資源配置,排程器會自動調整各階段的規模,在各種專用硬體加速器之間維持均衡的吞吐量。
4 Tokenizer
Tokenizer 是現代大規模模型的基礎建構單元。它以非監督方式學習一個瓶頸化的潛在空間 (Latent Space),把原始資料轉換為更有效率的表徵。具體來說,視覺 tokenizer 把原始且冗餘的視覺資料(如圖像與影片)映射為精簡的語意 token,對處理高維視覺資料至關重要。這種能力不僅讓大規模 transformer 模型得以高效訓練,也讓推論能在有限的計算資源上普及。圖 6 示意 token 化的訓練流程,目標是訓練編碼器與解碼器,使瓶頸處的 token 表徵最大程度保留輸入的視覺資訊。
圖 6:影片 token 化流程。 輸入影片被編碼為 token,通常比輸入影片精簡得多;解碼器再從 token 重建輸入影片。Tokenizer 的訓練就是學習編碼器與解碼器,讓 token 最大程度保留視覺資訊。【譯註:HTML 版未提供此示意圖圖檔。】


圖 7:連續與離散 tokenizer 的視覺化。 Token 沿空間($\frac{H}{s_{HW}}\times\frac{W}{s_{HW}}$)與時間($1+\frac{T}{s_{T}}$)維度排列,空間壓縮因子為 $s_{HW}$,時間壓縮因子為 $s_{T}$。第一個時間 token 代表第一個輸入影格,使圖像($T=0$)與影片($T>0$)能在共享潛在空間中聯合 token 化。左:嵌入維度為 $C$ 的連續潛在嵌入。右:量化索引,每種顏色代表一個離散潛在碼 (Latent Code)。
Tokenizer 分為連續與離散兩類(見圖 7)。連續 tokenizer 把視覺資料編碼為連續潛在嵌入,如 Stable Diffusion (Rombach et al. 2022) 或 VideoLDM (Blattmann et al. 2023b) 等潛在擴散模型 (Latent Diffusion Model) 所用。這類嵌入適合以連續分布取樣來生成資料的模型。離散 tokenizer 把視覺資料編碼為離散潛在碼,映射成量化索引 (Quantized Index),如 VideoPoet (Kondratyuk et al. 2024) 等自迴歸 transformer 所用。這種離散表徵對以交叉熵損失 (Cross-entropy Loss) 訓練的模型(如 GPT)是必要的。圖 7 描繪了這兩類 token。
Tokenizer 的成功很大程度取決於能否在不犧牲後續視覺重建品質的前提下提供高壓縮率。一方面,高壓縮能降低儲存與計算需求;另一方面,過度壓縮會導致關鍵視覺細節流失。這種權衡是 tokenizer 設計上的重大挑戰。
我們提出 Cosmos Tokenizer——一套視覺 tokenizer 套件,涵蓋圖像與影片的連續與離散 tokenizer。Cosmos Tokenizer 具備出色的視覺重建品質與推論效率,並提供多種壓縮率以因應不同的計算限制與應用需求。表 4 比較了不同視覺 tokenizer 及其能力。
| 模型 | 因果性 | 圖像 | 影片 | 聯合 | 離散 | 連續 |
|---|---|---|---|---|---|---|
| FLUX-Tokenizer (FLUX 2024) | - | ✓ | ✗ | ✗ | ✗ | ✓ |
| Open-MAGVIT2-Tokenizer (Luo et al. 2024) | - | ✓ | ✗ | ✗ | ✓ | ✗ |
| LlamaGen-Tokenizer (Sun et al. 2024a) | - | ✓ | ✗ | ✗ | ✓ | ✗ |
| VideoGPT-Tokenizer (Yan et al. 2021) | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ |
| Omni-Tokenizer (Wang et al. 2024a) | ✗ | ✓ | ✓ | ✓ | ✓ | ✓ |
| CogVideoX-Tokenizer (Yang et al. 2024d) | ✓ | ✓ | ✓ | ✓ | ✗ | ✓ |
| Cosmos-Tokenize1 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
表 4: 不同視覺 tokenizer 及其能力的比較。
我們以輕量、計算高效、具時間因果機制的架構設計 Cosmos Tokenizer。具體而言,我們採用因果時間卷積層 (Causal Temporal Convolution) 與因果時間注意力層,保留影片影格的自然時間順序,確保能以單一統一的網路架構無縫地 token 化圖像與影片。
我們直接在高解析度圖像與長時影片上訓練 tokenizer,不限制類別或長寬比。不同於聚焦特定資料類別與尺寸的既有 tokenizer,Cosmos Tokenizer 可處理各種長寬比——包括 1:1、3:4、4:3、9:16 與 16:9,且推論時對時間長度不設限,能 token 化超過訓練時所見時間長度的影片。
我們也在標準圖像與影片基準資料集上評估 tokenizer,包括 MS-COCO 2017 (Lin et al. 2014)、ImageNet-1K (Deng et al. 2009) 與 DAVIS (Perazzi et al. 2016)。為了推動物理 AI 應用的影片 token 化研究,我們整理了一個涵蓋多種物理 AI 影片類別的影片資料集,範圍包括魚眼鏡頭、機器人、駕駛、人類活動與空間導航。資料集見 github.com/NVlabs/TokenBench。
圖 8:連續(左)與離散(右)tokenizer 在時空壓縮率(對數尺度)與重建品質(PSNR)上的比較。 每個實心點代表一種 tokenizer 配置,呈現壓縮率與品質之間的權衡。值得注意的是,我們的 tokenizer 展現出色的壓縮-品質權衡,即使在更高壓縮率下品質仍優於其他方法。評估在 DAVIS 資料集上進行;圖像 tokenizer 的 PSNR 以逐影格方式計算。【譯註:此圖為互動式圖表,HTML 版未提供可下載之靜態圖檔。】
如圖 8 所示,評估結果顯示 Cosmos Tokenizer 以大幅優勢超越既有 tokenizer——例如在 DAVIS 影片上重建品質提升 +4 dB PSNR。它的速度最高快 $12\times$,且在單張 80GB 記憶體的 NVIDIA A100 GPU 上,可一次編碼最長 8 秒的 1080p 影片或 10 秒的 720p 影片而不會記憶體不足。
4.1 架構
Cosmos Tokenizer 採編碼器-解碼器架構。給定輸入影片 $x_{0:T}\in\mathbb{R}^{(1+T)\times H\times W\times 3}$($H$、$W$、$T$ 分別為高、寬與影格數),編碼器($\mathcal{E}$)把輸入 token 化為 token 影片 $z_{0:T^{\prime}}\in\mathbb{R}^{(1+T^{\prime})\times H^{\prime}\times W^{\prime}\times C}$,空間壓縮因子為 $s_{HW}=\frac{H}{H^{\prime}}=\frac{W}{W^{\prime}}$,時間壓縮因子為 $s_{T}=\frac{T}{T^{\prime}}$。解碼器($\mathcal{D}$)再從這些 token 重建輸入影片,得到重建影片 $\hat{x}_{0:T}\in\mathbb{R}^{(1+T)\times H\times W\times 3}$,數學式為:
$$\hat{x}_{0:T}=\mathcal{D}\Big(\mathcal{E}\big(x_{0:T}\big)\Big).\tag{1}$$
我們的架構採用時間因果設計,確保每一階段只處理當前與過去的影格,不依賴未來影格。與常見做法不同,我們的 tokenizer 在小波空間 (Wavelet Space) 中運作:輸入先經過 2 階小波轉換 (Wavelet Transform)。具體來說,小波轉換以分組方式對輸入影片 $x_{0:T}$ 沿 $x$、$y$、$t$ 各下採樣 4 倍。分組方式為:$\{x_{0},x_{1:4},x_{5:8},...,x_{(T-3):T}\}\rightarrow\{g_{0},g_{1},g_{2},...,g_{T/4}\}$。後續的編碼器階段以時間因果方式處理影格:$\{g_{0},g_{0:1},g_{0:2},...\}\rightarrow\{\xi_{0},\xi_{1},\xi_{2},...\}$。之後的編碼器階段遵循類似方案,最終輸出 token $z_{0:T^{\prime}}$。因果設計有助於建立在 tokenizer 之上的模型適配到通常運作於時間因果設定下的下游物理 AI 應用。小波轉換讓我們能在更精簡的影片表徵上運作,消除像素資訊的冗餘,使其餘各層專注於更語意層面的壓縮。
編碼器各階段(小波轉換之後)由一系列殘差區塊 (Residual Block) 與下採樣區塊交錯實作。在每個區塊中,我們採用時空分解的 3D 卷積:先用核尺寸 $1\times k\times k$ 的 2D 卷積擷取空間資訊,再用核尺寸 $k\times 1\times 1$ 的時間卷積擷取時間動態,並使用 $k-1$ 的左側填補 (Left Padding) 確保因果性。為了捕捉長程依賴,我們採用具全域支撐區域的時空分解因果自注意力 (Causal Self-attention)——例如最後一個編碼器區塊的支撐為 $1+T^{\prime}$。非線性使用 Swish 激活函數 (Ramachandran et al. 2017)。我們採用層正規化 (Layer Normalization, LayerNorm) (Lei Ba et al. 2016) 取代群組正規化 (Group Normalization, GroupNorm) (Wu and He 2018),避免潛在空間或重建輸出的特定區域出現過大的數值 (Karras et al. 2020; Sadat et al. 2024)。解碼器與編碼器鏡像對稱,把下採樣區塊換成上採樣區塊。圖 9 描繪 Cosmos Tokenizer 的整體架構。

(a)時間因果性: 時間因果機制示意,輸入 $x_{0},x_{1},\dots,x_{12}$ 經分組的中間輸出 $g_{0},g_{1},\dots$ 處理,再由時空卷積與注意力操作進一步精煉。

(b)網路架構: 編碼器-解碼器網路結構包含 3D Haar 小波、因果殘差、因果下採樣與因果時空注意力區塊。解碼器鏡像編碼器結構,把下採樣換成上採樣。
圖 9:Cosmos Tokenizer 整體架構,展示時間因果性與編碼器-解碼器結構的整合。 時間因果性(左)處理序列輸入,編碼器-解碼器(右)利用小波轉換與因果操作捕捉資料中的空間與時間依賴。
我們採用基本的自編碼器 (Autoencoder, AE) 形式來建模連續 tokenizer 的潛在空間。對離散 tokenizer,我們採用有限純量量化 (Finite-Scalar-Quantization, FSQ) (Mentzer et al. 2023) 作為潛在空間量化器。連續 tokenizer 的潛在維度為 $16$;離散 tokenizer 的潛在維度為 $6$,即 FSQ 層級數,層級為 $(8,8,8,5,5,5)$,對應詞彙表大小 $64{,}000$。
4.2 訓練策略
我們採用聯合訓練策略,以預設頻率交替使用圖像與影片小批次 (Mini-batch)。我們只對 tokenizer 解碼器的最終輸出施加監督,不使用作用於潛在空間的輔助損失,例如承諾損失 (Commitment Loss) 或 KL 先驗損失。舉例來說,若連續 tokenizer 採 VAE (Kingma 2013) 而非基本 AE,就需要 KL 先驗損失;若離散量化採 VQ-VAE (van den Oord et al. 2017) 而非 FSQ,就需要承諾損失。
我們採用兩階段訓練方案。第一階段以 L1 損失最佳化,最小化輸入與重建影片($\hat{x}_{0:T}$)之間的逐像素 RGB 差異:
$$\mathcal{L}_{1}=\left\lVert\hat{x}_{0:T}-x_{0:T}\right\rVert_{1},\tag{2}$$
以及基於 VGG-19 特徵 (Simonyan and Zisserman 2014) 的感知損失 (Perceptual Loss):
$$\mathcal{L}_{\text{Perceptual}}=\frac{1}{L}\sum_{l=1}^{L}\sum_{t}\alpha_{l}\left\lVert\texttt{VGG}_{l}(\hat{x}_{t})-\texttt{VGG}_{l}(x_{t})\right\rVert_{1},\tag{3}$$
其中 $\texttt{VGG}_{l}(\cdot)\in\mathbb{R}^{H\times W\times C}$ 是預訓練 VGG-19 網路第 $l$ 層的特徵,$L$ 是納入計算的層數,$\alpha_{l}$ 是第 $l$ 層的權重。

圖 10:TokenBench 的範例影片。 本圖展示多樣的範例,包括第一人稱視角、駕駛、機器人操作與網路影片。
第二階段使用光流損失 (Teed and Deng 2020) 處理重建影片的時間平滑度:
$$\mathcal{L}_{\text{Flow}}=\frac{1}{T}\sum_{t=1}^{T}\left\lVert\texttt{OF}(\hat{x}_{t},\hat{x}_{t-1})-\texttt{OF}(x_{t},x_{t-1})\right\rVert_{1}+\frac{1}{T}\sum_{t=0}^{T-1}\left\lVert\texttt{OF}(\hat{x}_{t},\hat{x}_{t+1})-\texttt{OF}(x_{t},x_{t+1})\right\rVert_{1},$$
並使用 Gram 矩陣損失 (Gatys et al. 2016) 增強重建圖像的銳利度:
$$\mathcal{L}_{\text{Gram}}=\frac{1}{L}\sum_{l=1}^{L}\sum_{t}\alpha_{l}\left\lVert\texttt{GM}_{l}(\hat{x}_{t})-\texttt{GM}_{l}(x_{t})\right\rVert_{1}.\tag{4}$$
此外,我們在微調階段使用對抗損失 (Adversarial Loss) 進一步強化重建細節,在大壓縮率下尤其有效。
我們以兩種壓縮率訓練圖像 tokenizer(記為 CI 與 DI):$8\times 8$ 與 $16\times 16$;以三種壓縮率訓練影片 tokenizer(記為 CV 與 DV):$4\times 8\times 8$、$8\times 8\times 8$ 與 $8\times 16\times 16$。此處壓縮率對圖像表示為 $H\times W$,對影片表示為 $T\times H\times W$,其中 $T$ 為時間維度,$H$ 與 $W$ 為空間維度。
影片 tokenizer 有兩個變體:
- Cosmos-0.1-Tokenizer: 以取樣較少 720p 影格的小批次訓練(CV 為 49 影格、DV 為 17 影格)。
- Cosmos-Tokenize1: 以取樣較多 360p 或 720p 影格的小批次訓練(CV 為 121 影格、DV 為 49 影格)。
這種做法確保能彈性處理圖像與影片資料的各種時間與空間解析度。實驗顯示,tokenizer 對訓練所見的解析度泛化良好,在更高解析度下也維持強勁品質。
4.3 結果
我們在多個圖像與影片基準資料集上廣泛評估 Cosmos Tokenizer 套件。圖像 tokenizer 的評估遵循先前做法,使用 MS-COCO 2017 (Lin et al. 2014) 與 ImageNet-1K (Deng et al. 2009):以 MS-COCO 2017 驗證子集的 $5{,}000$ 張圖像與 ImageNet-1K 驗證子集的 $50{,}000$ 張圖像作為圖像評估基準。
TokenBench。 影片 tokenizer 的評估目前還沒有針對高解析度、長時影片的標準基準。為此,我們推出 TokenBench 基準,涵蓋機器人操作、駕駛、第一人稱與網路影片等多種領域,並將評估標準化。我們取用各任務常用的既有影片資料集,包括 BDD100K (Yu et al. 2020)、EgoExo-4D (Grauman et al. 2024)、BridgeData V2 (Walke et al. 2023) 與 Panda-70M (Chen et al. 2024)。我們從每個資料集隨機取樣 $100$ 部影片,取前 $10$ 秒並把短邊縮放為 $1080$ 進行前處理。對 Panda-70M,我們人工過濾內容品質低與動作幅度小的影片。對 EgoExo-4D,我們隨機挑選 $100$ 個場景,各取樣一部第一人稱與一部第三人稱影片。最終共 $500$ 部影片。TokenBench 的部分範例見圖 10。我們於 github.com/NVlabs/TokenBench 釋出 TokenBench。
除了 TokenBench,我們也在 $1080$p 解析度的 DAVIS 資料集上評估影片 tokenizer。
基線與評估指標。 我們以多種壓縮率評估 tokenizer,展示其對不同計算需求的效益,並與最先進的圖像及影片 tokenizer 比較。表 4 列出各設定下比較的 SOTA tokenizer。評估指標包括峰值訊噪比 (Peak Signal-to-Noise Ratio, PSNR)、結構相似度 (Structural Similarity, SSIM)、圖像的重建 Fréchet Inception Distance (rFID) (Heusel et al. 2017),以及影片的重建 Fréchet Video Distance (rFVD) (Unterthiner et al. 2019)。
量化結果。 表 5 與表 6 匯總連續與離散影片 tokenizer 在各基準上的平均量化指標。如兩表所示,在 DAVIS 影片資料集與 TokenBench 上,Cosmos Tokenizer 在 $4\times 8\times 8$ 時空壓縮率下的所有指標皆達到最先進水準。此外,即使壓縮率高出 $2\times$ 與 $8\times$(即 $8\times 8\times 8$ 與 $8\times 16\times 16$),Cosmos Tokenizer 的品質仍優於先前方法,展現出色的壓縮-品質權衡。
| DAVIS | TokenBench | |||||||
|---|---|---|---|---|---|---|---|---|
| Tokenizer | 影格數 | 形式 | PSNR $\uparrow$ | SSIM $\uparrow$ | rFVD $\downarrow$ | PSNR $\uparrow$ | SSIM $\uparrow$ | rFVD $\downarrow$ |
| CogVideoX-Tokenizer4$\times$8$\times$8 | 17 | VAE | 29.29 | 0.864 | 19.58 | 32.06 | 0.909 | 6.97 |
| Omni-Tokenizer4$\times$8$\times$8 | 17 | VAE | 22.23 | 0.713 | 117.66 | 24.48 | 0.830 | 35.86 |
| Cosmos-0.1-Tokenizer-CV4$\times$8$\times$8 | 49 | AE | 32.80 | 0.900 | 15.93 | 35.45 | 0.928 | 6.85 |
| Cosmos-0.1-Tokenizer-CV8$\times$8$\times$8 | 49 | AE | 30.61 | 0.856 | 30.16 | 34.44 | 0.917 | 11.62 |
| Cosmos-0.1-Tokenizer-CV8$\times$16$\times$16 | 49 | AE | 27.60 | 0.779 | 93.82 | 31.61 | 0.875 | 43.08 |
| Cosmos-Tokenize1-CV4$\times$8$\times$8-360p | 49 | AE | 35.85 | 0.920 | 10.057 | 38.42 | 0.950 | 3.34 |
| Cosmos-Tokenize1-CV8$\times$8$\times$8-720p | 121 | AE | 31.28 | 0.868 | 23.49 | 35.13 | 0.926 | 9.82 |
表 5: 連續影片 (CV) tokenizer 在 DAVIS 與 TokenBench 上的評估。
| DAVIS | TokenBench | |||||||
|---|---|---|---|---|---|---|---|---|
| Tokenizer | 影格數 | 量化方式 | PSNR $\uparrow$ | SSIM $\uparrow$ | rFVD $\downarrow$ | PSNR $\uparrow$ | SSIM $\uparrow$ | rFVD $\downarrow$ |
| VideoGPT-Tokenizer4$\times$4$\times$4 | - | VQ | 28.17 | 0.850 | 72.33 | 33.66 | 0.914 | 13.85 |
| Omni-Tokenizer4$\times$8$\times$8 | 17 | VQ | 20.02 | 0.703 | 188.60 | 25.31 | 0.827 | 53.55 |
| Cosmos-0.1-Tokenizer-DV4$\times$8$\times$8 | 17 | FSQ | 28.81 | 0.818 | 37.36 | 31.97 | 0.888 | 19.67 |
| Cosmos-0.1-Tokenizer-DV8$\times$8$\times$8 | 17 | FSQ | 27.51 | 0.789 | 100.15 | 30.95 | 0.873 | 43.86 |
| Cosmos-0.1-Tokenizer-DV8$\times$16$\times$16 | 17 | FSQ | 25.09 | 0.714 | 241.52 | 28.91 | 0.829 | 113.48 |
| Cosmos-Tokenize1-DV4$\times$8$\times$8-360p | 49 | FSQ | 32.97 | 0.840 | 53.44 | 35.74 | 0.910 | 19.25 |
| Cosmos-Tokenize1-DV8$\times$16$\times$16-720p | 49 | FSQ | 25.49 | 0.719 | 259.33 | 29.33 | 0.838 | 107.43 |
表 6: 離散影片 (DV) tokenizer 在 DAVIS 與 TokenBench 上的評估。
表 7 與表 8 匯總連續與離散圖像 tokenizer 在多種圖像基準上的平均量化指標,涵蓋廣泛的圖像類型。如表所示,相較先前方法,Cosmos Tokenizer 在 $8\times 8$ 壓縮率下持續取得最先進的結果。更重要的是,在大 $4\times$ 的 $16\times 16$ 壓縮率下,Cosmos Tokenizer 的圖像品質常與先前方法在 $8\times 8$ 壓縮率下的品質相當甚至更好。
| (a) MS-COCO 2017 | (b) ImageNet-1K | |||||||
|---|---|---|---|---|---|---|---|---|
| Tokenizer | 高度 | 形式 | PSNR $\uparrow$ | SSIM $\uparrow$ | rFID $\downarrow$ | PSNR $\uparrow$ | SSIM $\uparrow$ | rFID $\downarrow$ |
| FLUX-Tokenizer8$\times$8 | - | VAE | 24.00 | 0.682 | 2.501 | 20.09 | 0.518 | 1.229 |
| Cosmos-0.1-Tokenizer-CI8$\times$8 | 1024 | AE | 28.66 | 0.836 | 1.760 | 28.83 | 0.837 | 0.689 |
| Cosmos-0.1-Tokenizer-CI16$\times$16 | 1024 | AE | 23.63 | 0.663 | 3.823 | 23.72 | 0.655 | 1.031 |
| Cosmos-Tokenize1-CI8$\times$8-360p | 1024 | AE | 32.79 | 0.824 | 1.874 | 32.91 | 0.824 | 0.785 |
| Cosmos-Tokenize1-CI16$\times$16-360p | 1024 | AE | 31.28 | 0.682 | 4.294 | 31.29 | 0.675 | 0.701 |
表 7: 連續圖像 (CI) tokenizer 在多種圖像資料集上的評估。
| (a) MS-COCO 2017 | (b) ImageNet-1K | |||||||
|---|---|---|---|---|---|---|---|---|
| Tokenizer | 高度 | 量化方式 | PSNR $\uparrow$ | SSIM $\uparrow$ | rFID $\downarrow$ | PSNR $\uparrow$ | SSIM $\uparrow$ | rFID $\downarrow$ |
| Open-MAGVIT2-Tokenizer16$\times$16 | - | LFQ | 19.50 | 0.502 | 6.649 | 17.00 | 0.398 | 2.701 |
| LlamaGen-Tokenizer8$\times$8 | - | VQ | 21.99 | 0.616 | 4.123 | 19.64 | 0.498 | 1.403 |
| LlamaGen-Tokenizer16$\times$16 | - | VQ | 19.11 | 0.491 | 6.077 | 18.38 | 0.448 | 1.657 |
| Cosmos-0.1-Tokenizer-DI8$\times$8 | 1024 | FSQ | 24.40 | 0.704 | 3.710 | 24.48 | 0.701 | 1.265 |
| Cosmos-0.1-Tokenizer-DI16$\times$16 | 1024 | FSQ | 20.45 | 0.529 | 7.234 | 20.49 | 0.518 | 2.518 |
| Cosmos-Tokenize1-DI8$\times$8-360p | 1024 | AE | 31.36 | 0.714 | 4.133 | 31.34 | 0.707 | 1.324 |
| Cosmos-Tokenize1-DI16$\times$16-360p | 1024 | AE | 30.54 | 0.565 | 7.557 | 30.46 | 0.549 | 2.147 |
表 8: 離散圖像 (DI) tokenizer 在多種圖像資料集上的評估。
這些在多種圖像與影片基準資料集上的量化結果證實,Cosmos Tokenizer 能在大幅時空壓縮下更好地表徵視覺內容。
執行效能。 表 9 列出參數量以及在單張 A100 80GB GPU 上量測的每張圖像或每一影格的平均編解碼時間,並列出先前最先進 tokenizer 的參數與平均速度作為比較。如表所示,無論圖像或影片 tokenizer,Cosmos Tokenizer 都快 $2\times\sim 12\times$,同時維持最小的模型規模,顯示其編解碼視覺內容的高效率。
| Tokenizer | 類型 | 解析度 | 影格數 | 參數量 | 時間 (ms) |
|---|---|---|---|---|---|
| FLUX-Tokenizer8$\times$8 | 連續-圖像 | $1024\times 1024$ | - | 84M | 242 |
| Cosmos-0.1-Tokenizer-CI8$\times$8 | 連續-圖像 | $1024\times 1024$ | - | 77M | 62.7 |
| LlamaGen-Tokenizer8$\times$8 | 離散-圖像 | $1024\times 1024$ | - | 70M | 475 |
| Cosmos-0.1-Tokenizer-DI8$\times$8 | 離散-圖像 | $1024\times 1024$ | - | 79M | 64.2 |
| CogVideoX-Tokenizer4$\times$8$\times$8 | 連續-影片 | $720\times 1280$ | 17 | 216M | 414 |
| Omni-Tokenizer4$\times$8$\times$8 | 連續-影片 | $720\times 1280$ | 17 | 54M | 82.9 |
| Cosmos-0.1-Tokenizer-CV4$\times$8$\times$8 | 連續-影片 | $720\times 1280$ | 49 | 105M | 34.8 |
| Omni-Tokenizer4$\times$8$\times$8 | 離散-影片 | $720\times 1280$ | 17 | 54M | 53.2 |
| Cosmos-0.1-Tokenizer-DV4$\times$8$\times$8 | 離散-影片 | $720\times 1280$ | 17 | 105M | 51.5 |
表 9: Tokenizer 執行效能比較。時間以每張圖像或每一影格為單位。
5 世界基礎模型預訓練
預訓練 WFM 是通才,捕捉真實世界物理與自然行為的一般知識。我們運用兩種可規模化的深度學習範式——擴散模型與自迴歸模型——建立兩個 WFM 家族。擴散模型與自迴歸模型都把困難的生成問題拆解為一系列較容易的子問題,一直是生成模型發展的強力推手。就擴散模型而言,困難的生成問題被拆解為一系列去噪 (Denoising) 問題;就自迴歸模型而言,則被拆解為一系列下一個 token 預測問題。我們將討論在建立預訓練 WFM 的過程中,如何運用為現代 GPU 量身打造的各種平行化技術來規模化這些深度學習範式。本文所有 WFM 模型都是在一座擁有 $10{,}000$ 張 NVIDIA H100 GPU 的叢集上、歷時三個月訓練而成。
| 類型 | 擴散 (Diffusion) | 自迴歸 (Autoregressive) |
|---|---|---|
| 模型 | 1. Cosmos-Predict1-7B-Text2World $\rightarrow$ Cosmos-Predict1-7B-Video2World 2. Cosmos-Predict1-14B-Text2World $\rightarrow$ Cosmos-Predict1-14B-Video2World |
1. Cosmos-Predict1-4B $\rightarrow$ Cosmos-Predict1-5B-Video2World 2. Cosmos-Predict1-12B $\rightarrow$ Cosmos-Predict1-13B-Video2World |
| Tokenizer | Cosmos-Tokenize1-CV8$\times$8$\times$8-720p | Cosmos-Tokenize1-DV8$\times$16$\times$16-720p |
| 強化器 | Cosmos-UpsamplePrompt1-12B-Text2World | Cosmos-Predict1-7B-Decoder-DV8$\times$16$\times$16ToCV8$\times$8$\times$8-720p |
表 10:Cosmos 世界基礎模型地圖。 我們有兩組 WFM:一組基於擴散模型,另一組基於自迴歸模型。每個家族各建立兩個基礎模型與兩個衍生模型。為了達到最佳生成品質,我們也為擴散模型建立提示詞上採樣器,為自迴歸模型建立擴散解碼器。
表 10 呈現我們的預訓練 WFM 及其配套模型的地圖。對基於擴散的 WFM 家族,我們先建立 7B 與 14B 兩個 Text2World 模型,即 Cosmos-Predict1-7B-Text2World 與 Cosmos-Predict1-14B-Text2World。這些模型能把文字提示詞映射為視覺世界的影片。接著我們微調 Text2World 模型,使其接受額外的影片輸入(代表當前觀察),得到 Video2World 模型——根據當前觀察(輸入影片)與擾動(文字提示詞)預測未來影片。這些擴散模型是接受連續 token 的潛在擴散模型,使用 Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 產生視覺 token。WFM 的訓練用文字提示詞由 VLM 透過影片描述生成產出,其分布與人類對影片的描述不同。為了縮小這個領域差距 (Domain Gap),我們基於 Mistral-NeMo-12B-Instruct 模型 (Mistral and NVIDIA 2024) 建立了 Cosmos-UpsamplePrompt1-12B-Text2World,協助把人類文字提示詞轉換為擴散 WFM 偏好的形式。
對基於自迴歸的 WFM 家族,我們先建立 4B 與 12B 兩個基礎模型,純粹根據當前影片觀察預測未來影片,分別命名為 Cosmos-Predict1-4B 與 Cosmos-Predict1-12B。這些是為影片預測任務從零訓練的 Llama3 風格 GPT 模型,不具語言理解能力。為了讓自迴歸 WFM 能利用文字資訊做下一個 token 預測,我們在 transformer 區塊中加入交叉注意力 (Cross-attention) 層,把輸入文字提示詞的 T5 嵌入引入 WFM。這些自迴歸 WFM 使用 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p,把輸入影片映射為少量整數。Tokenizer 的重度壓縮有時會導致不想要的失真。為了解決這個問題,我們透過微調 Cosmos-Predict1-7B-Text2World 模型建立擴散解碼器(Cosmos-Predict1-7B-Decoder-DV8$\times$16$\times$16ToCV8$\times$8$\times$8-720p),把 DV8$\times$16$\times$16 空間的離散 token 映射為 CV8$\times$8$\times$8 空間的連續 token。
5.1 基於擴散的世界基礎模型
我們的擴散 WFM 是在 tokenizer 學到的潛在空間中運作的潛在擴散模型,能以精簡、降維的方式表徵影片。這種設計選擇有多項優點:降低訓練與推論的計算成本,同時簡化去噪任務 (Rombach et al. 2022; Hoogeboom et al. 2024)。我們採用 Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 把影片 token 化為潛在表徵。
5.1.1 問題形式化
我們採用 EDM (Karras et al. 2022; Karras et al. 2024) 提出的方法訓練擴散 WFM。去噪器 (Denoiser) $D_{\theta}$ 在噪聲水準 $\sigma$ 下的去噪分數匹配 (Denoising Score Matching) 損失定義為:
$$\mathcal{L}(D_{\theta},\sigma)=\mathbb{E}_{\mathbf{x}_{0},\mathbf{n}}\Big[\big\lVert D_{\theta}(\mathbf{x}_{0}+\mathbf{n};\sigma)-\mathbf{x}_{0}\big\rVert^{2}_{2}\Big],\tag{5}$$
其中 $\mathbf{x}_{0}\sim p_{\rm{data}}$ 是從訓練集取樣的乾淨圖像或影片,$\mathbf{n}\sim\mathcal{N}\big(\mathbf{0},\sigma^{2}\mathbf{I}\big)$ 是獨立同分布 (i.i.d.) 的高斯雜訊,$D_{\theta}$ 是以噪聲水準為條件的神經網路,負責去噪受污染的樣本 $\mathbf{x}_{0}+\mathbf{n}$。我們遵循 EDM 提出的前置條件化 (Preconditioning) 設計來參數化 $D_{\theta}$。整體訓練損失定義為 $\mathcal{L}(D_{\theta};\sigma)$ 對噪聲水準的加權期望:
$$\mathcal{L}(D_{\theta})=\mathbb{E}_{\sigma}\left[\frac{\lambda(\sigma)}{e^{u(\sigma)}}\mathcal{L}(D_{\theta},\sigma)+u(\sigma)\right],\tag{6}$$
$$\lambda(\sigma)=\big(\sigma^{2}+\sigma_{\text{data}}^{2}\big)\,/\,(\sigma\cdot\sigma_{\text{data}})^{2},\tag{7}$$
$$\ln(\sigma)\sim\mathcal{N}\big(P_{\text{mean}},P_{\text{std}}^{2}\big),\tag{8}$$
其中噪聲水準 $\sigma$ 的分布由超參數 $P_{\text{mean}}$ 與 $P_{\text{std}}$ 控制。$\sigma_{\text{data}}$ 是訓練資料的標準差,加權函數 $\lambda(\sigma)$ 確保訓練初期各噪聲水準的貢獻相等。然而隨著訓練推進,這種平衡可能惡化。為了緩解此問題,我們把對各噪聲水準的最佳化視為一種多任務學習 (Multi-task Learning):採用基於不確定性的加權方法,引入 $u(\sigma)$ 作為連續的不確定性函數,量化噪聲水準 $\sigma$ 下去噪目標 $\mathcal{L}(D_{\theta},\sigma)$ 的不確定性。我們用一個簡單的 MLP 參數化 $u(\sigma)$,並在訓練中最小化整體損失 $\mathcal{L}(D_{\theta})$。直觀而言,若模型對某任務不確定(即 $u(\sigma)$ 高),該噪聲水準的損失貢獻會被調降;同時模型會因這種不確定性受到懲罰,促使 $u(\sigma)$ 盡可能低。
相較於近期採用高斯流匹配 (Gaussian Flow Matching) 形式的影片生成模型 (Polyak et al. 2024; Kong et al. 2024),我們的工作源自擴散分數匹配觀點 (Ho et al. 2020; Song et al. 2020)。不過如 Gao et al. 2024a 所示,這些框架在理論上等價,其目標與訓練程序具有根本上的相似性。我們基於 EDM 的形式化與這些洞見一致,主要差別在於前置條件化設計與超參數的選擇。實務上,我們未遇到 EDM 形式化帶來的任何效能限制。
5.1.2 架構
本節描述去噪器網路 $D_{\theta}$ 的設計。它建立在 DiT (Peebles and Xie 2023) 之上——DiT 原本是為標籤條件式圖像生成設計的,我們調整其架構使其更適合可控影片生成的目標。整體網路設計見圖 11。

圖 11:Cosmos-Predict1 世界基礎模型整體架構。 模型把輸入影片經 Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 的編碼器處理得到潛在表徵,隨後加入高斯雜訊擾動。這些表徵再經 3D patch 化 (Patchification) 轉換。在潛在空間中,架構重複套用自注意力、交叉注意力(整合輸入文字)與前饋 MLP 層構成的區塊,並由對應時間步 $t$ 的自適應層正規化(縮放、平移、閘控)調變。Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 的解碼器再從精煉後的潛在表徵重建最終影片輸出。
3D patch 化。 網路的輸入是形狀為 $T\times C\times H\times W$ 的潛在表徵(圖像與影片皆然,圖像視為單一影格的影片)。為了替去噪器網路準備輸入,我們先用線性層把狀態「patch 化」再攤平。此過程把形狀 $(p_{t},p_{h},p_{w})$ 的不重疊立方體投影為網路的個別 token 輸入。因此,patch 化後的圖像或影片被重塑為長度 $THW/(p_{t}p_{h}p_{w})$ 的一維時空序列。我們的去噪器網路使用 $p_{t}=1,p_{h}=p_{w}=2$。
FPS 感知 3D RoPE 與可學習嵌入的混合位置嵌入。 我們採用 3D 分解的旋轉位置嵌入 (Rotary Position Embedding, RoPE) (Su et al. 2024),以支援任意尺寸、長寬比與影片長度的生成。具體來說,我們把特徵維度劃分為三個大致相等的區塊,分別沿時間、高度、寬度軸套用帶位置資訊的 RoPE。實務上,這可以透過在各自軸上串接頻率嵌入並重用為 LLM 最佳化的 RoPE kernel 來高效實作,不需在每個區塊中切分與串接。為了進一步支援不同影格率 (Frames Per Second, FPS) 的影片合成,我們根據訓練影片的 FPS 重新縮放時間頻率。由於 RoPE 的相對位置編碼特性與我們的 3D 分解設計,FPS 感知設計與圖像-影片聯合訓練相容。RoPE 的另一個好處在漸進式訓練中改變解析度或影片長度時顯現:藉助 Neural Tangent Kernel (NTK)-RoPE (Peng and Quesnelle 2023),我們觀察到模型快速收斂,即使在 $5{,}000$ 步訓練內也能達到合理表現。此外,我們發現每個 transformer 區塊額外加入一個可學習的絕對位置嵌入能進一步強化模型、降低訓練損失,並減少生成影片中的形變瑕疵 (Morphing Artifact)。
用於文字條件化的交叉注意力。 我們依靠網路中的交叉注意力層引入語言資訊。每個 transformer 區塊由依序排列的自注意力、交叉注意力與前饋層組成。自注意力作用於時空 token,交叉注意力則以 T5-XXL (Raffel et al. 2020) 嵌入作為 key 與 value 整合語意脈絡,實現有效的文字條件化。
Query-Key 正規化。 在訓練初期,我們觀察到注意力 logits 增長不穩定,導致注意力熵崩塌 (Attention Entropy Collapse)。我們遵循既有文獻 (Dehghani et al. 2023; Wortsman et al. 2023; Esser et al. 2024),在注意力運算前對 query $Q$ 與 key $K$ 做正規化。我們對網路中所有自注意力與交叉注意力層使用帶可學習縮放的均方根正規化 (Root Mean Square Normalization, RMSNorm) (Zhang and Sennrich 2019)。
AdaLN-LoRA。 我們發現 DiT 的自適應層正規化 (Adaptive Layer Normalization, AdaLN) 層 (Xu et al. 2019; Peebles and Xie 2023) 佔用大量模型參數,但以 FLOPs 而言對計算複雜度的貢獻微乎其微。受 W.A.L.T (Gupta et al. 2024a) 啟發,我們實作低秩適應 (Low-Rank Adaptation, LoRA) (Hu et al. 2022),把這些層中的稠密線性投影分解為低秩近似。對 Cosmos-Predict1-7B,這項架構最佳化把參數量減少 36%(從 11B 降到 7B),同時在所有評估指標上維持相同表現,證明了我們參數高效設計的有效性。
| 配置 | 7B-Text2World | 14B-Text2World | 7B-Video2World | 14B-Video2World |
|---|---|---|---|---|
| 層數 | $28$ | $36$ | $28$ | $36$ |
| 模型維度 | $4{,}096$ | $5{,}120$ | $4{,}096$ | $5{,}120$ |
| FFN 隱藏維度 | $16{,}384$ | $20{,}480$ | $16{,}384$ | $20{,}480$ |
| AdaLN-LoRA 維度 | $256$ | $256$ | $256$ | $256$ |
| 注意力頭數 | $32$ | $40$ | $32$ | $40$ |
| Key / Value 頭數 | $32$ | $40$ | $32$ | $40$ |
| MLP 激活函數 | GELU | |||
| 位置嵌入 | 混合位置嵌入 | |||
| 條件資訊 | 文字;FPS | 文字;FPS | 文字;FPS;影格 | 文字;FPS;影格 |
| 基礎學習率 | $2^{-15}$ | $2^{-16}$ | $2^{-15}$ | $2^{-16}$ |
| 權重衰減 | $0.1$ | $0.2$ | $0.1$ | $0.2$ |
| 學習率預熱 | 線性排程,$2{,}500$ 次迭代 | |||
| AdamW 動量與 $\epsilon$ | $\beta_{1},\beta_{2}=0.9,0.99$;$\epsilon=10^{-10}$ |
表 11: Cosmos-Predict1 模型的配置細節。
5.1.3 訓練策略
本節概述我們在涵蓋多種模態、解析度、長寬比與條件輸入的資料集上訓練模型的方法。
圖像與影片聯合訓練。 為了在模型訓練中利用大量高品質、多樣化的圖像資料集,我們實作交替最佳化策略,交錯處理圖像與影片批次。為了促進圖像與影片領域間的跨模態知識轉移,我們採用領域專屬的正規化方案,使用對圖像與影片資料分別估計的充分統計量 (Sufficient Statistics) 對齊潛在分布。此做法的動機是我們觀察到,縮小圖像與影片潛在表徵之間的分布偏移能提升生成品質。此外,我們觀察到影片潛在表徵在時間與通道維度上的統計量並不平穩 (Non-stationary)。為了處理這種異質性,我們採用逐影格、逐通道標準化的正規化策略,有效促使影片潛在表徵更接近等向性高斯先驗分布 (Isotropic Gaussian Prior)。
除了跨模態知識轉移,我們的正規化方案還帶來一項重要的理論好處:訓練期間訊噪比的尺度不變性 (Scale Invariance)。考慮兩個零均值但尺度不同的潛在表徵:一個標準化為單位變異數,另一個變異數為 4。當對標準化表徵加入高斯雜訊 $\mathcal{N}(0,\sigma^{2})$ 以達到期望的訊噪比時,對未正規化的表徵就必須把雜訊放大為 $\mathcal{N}(0,4\sigma^{2})$ 才能維持相同比率。把所有潛在表徵標準化後,我們確保不同尺度下訊噪比的一致性,即使訓練期間更新了底層 tokenizer,模型也能順利適應。
為了維持計算效率,我們平衡圖像與影片的批次大小,確保各 GPU 的記憶體使用相當。然而我們觀察到影片批次的去噪損失收斂比圖像批次慢。我們將其歸因於影片影格固有的時間冗餘,導致影片批次的梯度幅度較小。借鏡多解析度圖像訓練的近期進展 (Chen 2023; Hoogeboom et al. 2023; Atzmon et al. 2024),我們把影片批次的噪聲水準相對於圖像批次縮放影格數的平方根,以解決這種收斂差異。
| 階段 | 解析度 | 影格數 | 上下文長度 | FSDP 大小 | CP 大小 |
|---|---|---|---|---|---|
| 低解析度預訓練 | 512p (640$\times$512) | 57 | $10{,}240$ ^a | 64 | 2 |
| 高解析度預訓練 | 720p (1280$\times$704) | 121 | $56{,}320$ ^b | 64 | 8 |
| 高品質微調 | 720p (1280$\times$704) | 121 | $56{,}320$ ^b | 64 | 8 |
表 12: 漸進式訓練各階段及其規格。
漸進式訓練。 我們採用漸進式訓練策略,各階段細節見表 12。第一階段在 512 像素解析度的影片與圖像上訓練,影片為 57 影格;隨後過渡到目標解析度 720 像素,影片長度增加到 121 影格。在大量資料上預訓練後,我們在高品質子集上以線性衰減的學習率微調 $\mathcal{O}(10k)$ 次迭代。與 Dai et al. 2023 的發現一致,我們也發現微調能提升生成影片的品質。
多長寬比訓練。 為了容納不同長寬比的內容,我們把資料組織為五個桶 (Bucket),對應 1:1、3:4、4:3、9:16 與 16:9,並把每張圖像或每部影片分配到長寬比最接近的桶。訓練時,每個資料平行處理群組從一個桶取樣,不同平行處理群組可使用不同的桶。我們採用最長邊縮放 (Longest-side Resizing),最大程度保留提示詞所描述的原始內容資訊。批次處理時,我們對缺失像素施加鏡射填補 (Reflection Padding),並把填補遮罩提供給擴散主幹網路,以便推論時精確控制。
混合精度訓練。 我們保有兩份模型權重:一份 BF16、一份 FP32。前向與反向傳播使用 BF16 權重以提升訓練效率,梯度與激活值也因此為 BF16 格式。參數更新則以 FP32 進行以確保數值穩定,更新後的 FP32 參數再複製並轉型為 BF16 供下一次迭代使用。為了進一步穩定訓練,我們把式 (5) 的去噪分數匹配損失放大 10 倍。我們也發現 AdamW 使用較低的 beta 與 eps 係數能顯著減少損失尖峰 (Loss Spike)。在 14B 擴散模型的訓練中,我們極少遇到損失尖峰,且沒有出現不可恢復的損失尖峰。
文字條件化。 Text2World 模型採用 T5-XXL (Raffel et al. 2020) 作為文字編碼器。我們對 T5 嵌入做零填補以維持固定序列長度 512。為了強化文字與內容的對齊,我們採用無分類器引導 (Classifier-free Guidance) (Ho and Salimans 2022)。與先前隨機把文字嵌入歸零的工作 (Balaji et al. 2022; Saharia et al. 2022) 不同,我們省略此步驟,因為推論時的負向提示詞 (Negative Prompt) 已足夠有效。值得注意的是,作為文字生成圖像的模型,我們的模型即使不用引導也能生成高擬真圖像,我們把這歸功於高品質的訓練資料集。無分類器引導通常會促成偏好視覺內容的尋峰 (Mode-seeking) 行為,而我們發現謹慎的資料挑選能達到類似效果。然而在影片生成方面,缺乏同等高品質的資料使低引導設定下的結果欠佳,因此影片生成任務需要較高的引導值才能產生令人滿意的內容。
圖像與影片條件化。 我們擴充 Text2World 模型,建立支援圖像與影片條件化的 Video2World 模型,把先前影格納入生成過程。具體來說,條件影格與生成影格沿時間維度串接。為了提升推論時對輸入影格變異的穩健性,我們在訓練時對條件影格加入增強雜訊 (Augmented Noise),其 sigma 值以 $P_{\text{mean}}=-3.0,P_{\text{std}}=2.0$ 取樣。此外,擴散模型的輸入沿通道維度串接一個二元遮罩,用來區分條件影格與生成影格。損失函數排除條件影格位置的貢獻,只關注生成的輸出。為了提升泛化能力,訓練時我們隨機改變條件影格數。推論時,模型可以彈性地以單一條件影格(圖像)或多個先前影格作為輸入。
5.1.4 規模化
本節概述讓擴散 WFM 高效規模化的技術。我們分析模型的記憶體需求、討論平行化策略,並把我們的訓練配置與其他影片擴散模型及最先進的 LLM 做比較。
| 層 | 操作 | FLOPs | 激活值(張量形狀) |
|---|---|---|---|
| 自注意力 | $Q,K,V$ 投影 | $2\times 3\times\text{seq\_len}\times\text{d\_model}^{2}$ | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ ^a |
| $QK$ Norm | — | $2\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ ^b | |
| $A=Q@K^{T}$ | $2\times\text{seq\_len}^{2}\times\text{d\_model}$ | — ^c | |
| $A^{\prime}=\text{Softmax}(A)$ | — | — ^d | |
| $A^{\prime}@V$ | $2\times\text{seq\_len}^{2}\times\text{d\_model}$ | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ ^e | |
| 最終投影 | $2\times\text{seq\_len}\times\text{d\_model}^{2}$ | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ | |
| 交叉注意力 | $Q,K,V$ 投影 | $2\times\text{seq\_len}\times\text{d\_model}^{2}$ | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ |
| $QK$ Norm | — | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ ^f | |
| $A=Q@K^{T}$ | — | — ^c | |
| $A^{\prime}=\text{Softmax}(A)$ | — | — ^d | |
| $A^{\prime}@V$ | — | — ^g | |
| 最終投影 | $2\times\text{seq\_len}\times\text{d\_model}^{2}$ | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ | |
| 前饋層 | 上投影 | $4\times\text{seq\_len}\times\text{d\_model}^{2}$ | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ |
| GELU | — | $4\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ | |
| 下投影 | $4\times\text{seq\_len}\times\text{d\_model}^{2}$ | — ^h | |
| AdaLN | LayerNorm | — | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ |
| Scale | — | — ^i | |
| Shift | — | — | |
| Gate | — | $\text{seq\_len}\times\text{batch\_size}\times\text{d\_model}$ |
表 13:Cosmos-Diffusion transformer 的 FLOPs 與激活記憶體。 本表列出各操作的計算成本 (FLOPs) 與激活記憶體需求。FLOPs 以係數 2 描述乘加成本。「—」表示該值因數量級太小可忽略,或因採用激活檢查點 (Activation Checkpointing) 以重算取代儲存來節省記憶體而省略。
記憶體需求。 消耗 GPU 記憶體的四大元件為:
- 模型參數: 每參數 10 位元組。混合精度訓練同時以 FP32 與 BF16 儲存模型參數,另有 FP32 的指數移動平均 (Exponential Moving Average, EMA) 權重。
- 梯度: 每參數 2 位元組。梯度以 BF16 儲存。
- 最佳化器狀態: 每參數 8 位元組。我們使用 AdamW (Loshchilov and Hutter 2019) 作為最佳化器,其狀態(即一階與二階動差)以 FP32 儲存。
- 激活值: $(2\times\text{number\_of\_layers}\times 15\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model})$ 位元組,以 BF16 儲存。表 13 詳列網路中主要操作所儲存的激活值。為了最佳化記憶體使用,我們實作選擇性激活檢查點 (Chen et al. 2016; Korthikanti et al. 2023),對正規化函數等記憶體受限的層重算激活值。
舉例來說,14B 模型(Cosmos-Predict1-14B-Text2World)在高解析度預訓練時,模型參數、梯度與最佳化器狀態約需 280 GB,激活值另需 310 GB。鑑於 NVIDIA H100 GPU 的 80GB HBM3 上限,我們採用全分片資料平行 (Fully Sharded Data Parallelism, FSDP) 與上下文平行 (Context Parallelism, CP) 把記憶體需求分散到多張 GPU。
全分片資料平行 (FSDP)。 FSDP 把模型參數、梯度與最佳化器狀態分片到各裝置以提升記憶體效率:只在計算需要時聚集參數,用完即釋放。不同於在各裝置複製參數的標準資料平行,FSDP 把參數、梯度與最佳化器狀態分散開來,每個裝置只管理自己的分片。這種方法把記憶體使用降到「最大的暫時未分片參數集」加上自身分片的參數、梯度與最佳化器狀態。我們的實作對 7B 模型使用分片因子 32、對 14B 模型使用 64,以平衡記憶體與通訊延遲。
上下文平行 (CP)。 在長上下文設定下擴展 transformer 會帶來 FLOPs 與激活記憶體增加的挑戰。CP 透過把計算與激活值分散到多張 GPU 來解決:把 query $Q$ 與 key-value $(K,V)$ 沿序列維度切成 CP_SIZE 塊(CP_SIZE 為 CP 群組內的 GPU 數)。每張 GPU 處理一塊 $Q$,並反覆使用同一 CP 群組中儲存的 $(K,V)$ 區塊累積部分注意力輸出。CP 的不同實作使用不同通訊原語,包括 all-gather (Dubey et al. 2024)、P2P (Liu et al. 2023a) 與 all-to-all (Jacobs et al. 2023)。我們採用 TransformerEngine (NVIDIA 2024e) 的 P2P 變體,在 GPU 間傳輸 $(K,V)$ 區塊的同時處理注意力,使計算與通訊重疊。若區塊大小選擇得當,這種重疊能有效隱藏資料傳輸延遲。我們把 CP 群組組織在 NVLink 相連的 GPU 內,並讓 CP rank 與 FSDP rank 重疊以達最佳利用率。對上下文較短的圖像迭代,停用 CP 以提升吞吐量。交叉注意力層不使用 CP,因為其 $(K,V)$ 序列較短,計算量不足以掩蓋通訊延遲。
以 Cosmos-Predict1-14B 為例,採用分片因子 64 的 FSDP 後,參數、梯度與最佳化器狀態的記憶體需求從 280 GB 降到約 $280\mathbin{/}64\approx 4\ \text{GB/GPU}$;同樣地,採用 $\text{CP\_SIZE}=8$ 的 CP 後,激活記憶體從 310 GB 降到約 $310\mathbin{/}8\approx 40\ \text{GB/GPU}$。要注意這些計算是低估值:實務上 tokenizer 與未分片參數會消耗額外記憶體,CP 中通訊與計算的重疊也要求每張 GPU 保留多塊 $(K,V)$。
與其他影片生成模型的比較。 相較於 HunyuanVideo (Kong et al. 2024) 與 MovieGen (Polyak et al. 2024) 採用張量平行 (Tensor Parallelism, TP) 及其延伸序列平行 (Sequence Parallelism, SP),我們的平行化策略刻意保持精簡。儘管不用 TP/SP,我們的配置仍達到相當的模型 FLOPs 利用率 (Model FLOPs Utilization, MFU)。TP/SP 在某些情境(如更大的模型或不同的網路拓撲)仍有價值,詳細的權衡分析留待未來工作。
| 影格 0 | 影格 29 | 影格 59 | 影格 89 | 影格 120 | |
|---|---|---|---|---|---|
| 7B | ![]() |
![]() |
![]() |
![]() |
![]() |
| 14B | ![]() |
![]() |
![]() |
![]() |
![]() |
提示詞(節錄翻譯):雙手穩穩握住蒸氣熨斗的把手,熟練地在皺褶的襯衫上滑動。每一次來回,熨斗釋放出柔和的蒸氣,輕鬆撫平布料、消除皺痕,呈現平整俐落的效果。熨斗精準而細心地移動,一筆一劃地讓襯衫煥然一新。空氣中瀰漫著淡淡的清新亞麻香氣,柔和的光線從一旁的窗戶灑入,突顯布料新近撫平的質地,營造出寧靜的氛圍。
圖 12:Cosmos-Predict1-7B-Text2World 與 Cosmos-Predict1-14B-Text2World 生成的影片。 兩個 Text2World 模型都能產生視覺品質高、動態豐富且與文字對齊的影片。值得注意的是,相較 7B 模型,14B 模型展現出捕捉更精細視覺細節與更複雜運動模式的能力。完整影片與更多範例請造訪我們的網站。
| 條件影格 0 | 影格 29 | 影格 59 | 影格 89 | 影格 120 | |
|---|---|---|---|---|---|
| 7B | ![]() |
![]() |
![]() |
![]() |
![]() |
| 14B | ![]() |
![]() |
![]() |
![]() |
![]() |
提示詞(節錄翻譯):影片描繪一支機械手臂握著盛有紅酒的酒杯。這支配備多個關節與機械元件的手臂顯然是為精密任務設計。酒杯被輕柔地握住,展現機器人處理易碎物品的能力。背景極簡,突顯機器人與酒杯之間的互動。
| 條件影格 0 | 影格 150 | 影格 300 | 影格 450 | 影格 680 | |
|---|---|---|---|---|---|
| 7B | ![]() |
![]() |
![]() |
![]() |
![]() |
| 14B | ![]() |
![]() |
![]() |
![]() |
![]() |
提示詞(節錄翻譯):影片描繪一座大型工業設施的內部,可能是工廠或倉庫。空間寬敞、天花板挑高、金屬構架分明。可見高架起重機與各種機具,顯示這是重型製造或組裝場所。地面大致淨空,散落些許雜物並有標線。現場設有安全標誌與圍欄,強調其工業環境。自然光從高窗灑入,照亮工作區域。
圖 13:Cosmos-Predict1-7B-Video2World 與 Cosmos-Predict1-14B-Video2World 生成的影片。 上方兩列為以前 9 個影格為條件生成的 5 秒影片;下方兩列為長影片生成結果。我們以自迴歸方式生成長影片:第一段影片以單張輸入圖像為條件,其後五段各以前一段的最後 9 個影格為條件。7B 與 14B 模型都能產生高視覺擬真度的逼真影片;14B 模型展現生成更複雜場景的能力與更佳的運動穩定性。完整影片與更多範例請造訪我們的網站。
與大型語言模型的比較。 不同於通常以較短上下文長度預訓練的 LLM,長上下文設定因自注意力的二次方成本而顯著增加 FLOPs。LLM 的 FLOPs 常以 $6\times\text{seq\_len}\times P$ 計算($P$ 為參數量)(Kaplan et al. 2020),但我們指出這個公式對我們的擴散 WFM 並不準確。各關鍵操作的前向傳播 FLOPs 見表 13。
5.1.5 提示詞上採樣器
訓練期間,我們的 WFM 以詳細的影片描述作為輸入文字提示詞來產生高品質影片。然而推論時,使用者提示詞的長度、結構與風格各異,通常短得多。為了彌合訓練與推論提示詞之間的差距,我們開發了提示詞上採樣器,把原始輸入提示詞轉換為更詳細、更豐富的版本。它能為提示詞補充更多細節並維持一致的描述結構,帶來更高品質的輸出。
提示詞上採樣器的主要要求為:
- 對輸入提示詞的忠實度: 上採樣後的提示詞必須忠實保留原始使用者輸入的關鍵元素,包括主要角色、動作或運動、關鍵屬性與整體意圖。
- 與訓練分布對齊: 上採樣後的提示詞在長度、語言結構與風格上應貼近 WFM 訓練提示詞的分布。
- 強化視覺細節: 上採樣後的提示詞應能引導 WFM 生成更準確的影像。
Text2World 模型的提示詞上採樣器。 我們微調 Mistral-NeMo-12B-Instruct (Mistral and NVIDIA 2024) 建立提示詞上採樣器。為了取得配對資料——模擬使用者輸入的短提示詞與反映訓練提示詞分布的長提示詞——我們用 VLM 根據訓練用長提示詞與對應影片生成短描述。這種「由長生短」的資料建立策略能有效 (1) 保留 WFM 詳細訓練提示詞中的真實影片內容與分布,(2) 確保短提示詞與長提示詞之間的忠實度。所得的提示詞上採樣器命名為 Cosmos-UpsamplePrompt1-12B-Text2World。
Video2World 模型的提示詞上採樣器。 Video2World 模型的輸入包含影片條件與使用者文字提示詞。為了強化使用者提示詞,我們使用開源 VLM Pixtral-12B (Agrawal et al. 2024) 搭配零樣本提示工程 (Zero-shot Prompt Engineering),把提示詞上採樣為同時考量影片條件與使用者提示詞的詳細描述。我們發現原始的 Pixtral-12B 模型開箱即用效果良好,因此沒有進行前述的類似微調。
5.1.6 結果
圖 12 呈現 Cosmos-Predict1-7B-Text2World 與 Cosmos-Predict1-14B-Text2World 模型生成的質性結果。兩個模型都能產生視覺品質高、動態豐富且與文字對齊的影片。相較 7B 模型,14B 模型能生成捕捉更複雜視覺細節與更精緻運動的影片。
圖 13 展示 Video2World 7B 與 14B 模型生成的影片。Video2World 模型支援圖像與影片條件化,並能以自迴歸方式生成更長的影片。如圖 13 所示,我們的 Video2World 模型能產生動態良好、視覺擬真度高的逼真影片。14B 模型在場景豐富度與運動穩定性上同樣表現更佳。
5.2 基於自迴歸的世界基礎模型
在自迴歸 WFM 中,我們把世界模擬生成表述為類似語言建模的下一個 token 預測任務。我們先用第 4 節介紹的 Cosmos 離散 tokenizer 把影片轉換為離散影片 token 序列 $\mathcal{V}=\{v_{1},v_{2},\dots,v_{n}\}$,再訓練一個 Transformer 解碼器 (Vaswani et al. 2017),以過去的影片 token 為上下文預測下一個影片 token,做法與大型語言模型 (Brown et al. 2020; Jiang et al. 2023; Dubey et al. 2024) 類似。具體來說,訓練目標是最小化以下負對數概似 (Negative Log-likelihood, NLL) 損失:
$$\mathcal{L}_{NLL}=\sum_{i}-\log P(v_{i}|v_{1},v_{2},\dots,v_{i-1};\Theta),\tag{9}$$
其中預測下一個影片 token $v_{i}$ 的條件機率 $P$ 由參數為 $\Theta$ 的 Transformer 解碼器建模。
5.2.1 架構
我們的自迴歸 WFM 架構如圖 14 所示。我們針對影片生成任務對標準 transformer 模型架構做了幾項修改,包括加入 1) 3D 感知的位置嵌入、2) 交叉注意力以支援文字輸入、提升可控性,以及 3) QK 正規化 (Wortsman et al. 2023)。

圖 14:Cosmos-Predict1-Video2World 模型架構。 流程先把輸入影片經 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p 的編碼器編碼為離散 token,再轉換為可學習的嵌入。這些嵌入經過重複的 transformer 區塊處理,每個區塊包含絕對位置嵌入與 3D RoPE 元件,攤平後進入自注意力模組。每個區塊還包含一個交叉注意力模組,引入(經 T5 文字編碼器處理的)編碼文字提示詞,其後接兩層 MLP。最後由 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p 的解碼器從輸出 token 重建影片。
3D 位置嵌入。 與擴散 WFM(5.1.2 節)類似,我們納入兩種互補的位置嵌入機制:表達相對位置的 3D 分解 RoPE,以及表達絕對座標的 3D 分解絕對位置嵌入 (Absolute Positional Embedding, APE)。兩者協同運作,在整個網路中提供完整的空間與時間資訊。
- 3D 旋轉位置嵌入 (RoPE)。 我們對模型套用 3D RoPE,沿時間、高度、寬度維度編碼相對位置資訊。訓練時我們採用多階段策略,影片序列長度隨訓練推進而增加。為了讓 3D RoPE 適應變化的時間長度,我們使用 YaRN (Peng et al. 2023)——一種為延伸 RoPE 上下文視窗設計的高計算效率技術。由於影片序列長度只沿時間維度增加,我們只在時間軸套用 YaRN 延伸。藉助 YaRN,模型能外推到比訓練初期更長的上下文長度。
- 3D 絕對位置嵌入 (APE)。 除了 3D RoPE,我們在每個 transformer 區塊中加入 3D APE 以補足相對位置編碼。此 APE 使用沿時間、高度、寬度維度分解的正弦嵌入編碼位置資訊,確保模型知曉絕對位置。嵌入在每個階段直接加到輸入張量上,豐富 transformer 的位置脈絡。我們發現結合絕對與相對位置編碼能提升模型表現、降低訓練損失,並將生成影片中的形變瑕疵減到最少。值得注意的是,擴散 WFM(5.1.2 節)採用可學習嵌入,而自迴歸 WFM 的 APE 採用正弦式嵌入。
詞彙表。 在大型語言模型中,token 化是把輸入文字轉為離散 token 序列的關鍵步驟。LLM 的可能 token 詞彙表由其 tokenizer 決定(例如 OpenAI 2022 推出的 tiktoken),這些 tokenizer 以位元組對編碼 (Byte Pair Encoding, BPE) (Gage 1994) 等演算法在大型文字語料上訓練。
我們的自迴歸模型使用 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p 作為 tokenizer。如第 4 節所述,我們利用有限純量量化 (FSQ) (Mentzer et al. 2023) 把 $6$ 維潛在空間量化為 $(8,8,8,5,5,5)$ 個層級,對應詞彙表大小 $8\times 8\times 8\times 5\times 5\times 5=64{,}000$。
用於文字條件化的交叉注意力。 除了 transformer 架構中原有的自注意力區塊,我們加入交叉注意力層讓模型能以輸入文字為條件。與擴散 WFM(5.1.2 節)類似,交叉注意力作用於 transformer 模型的特徵與預訓練文字編碼器 (T5-XXL) 產生的文字嵌入之間。在我們的實驗中,每個自注意力層之後都加入交叉注意力區塊。
Query-Key 正規化。 為了增強訓練穩定性,我們納入 QKNorm (Wortsman et al. 2023)。QKNorm 在計算內積前先正規化 query($Q$)與 key($K$)向量,防止 softmax 函數飽和、確保更有效的學習,從而解決注意力機制的不穩定問題。正規化後,內積以可學習參數 $\gamma$ 縮放,而非固定的 $1/\sqrt{d_{k}}$。這個可學習的縮放因子讓模型能自適應地控制注意力分數的幅度,增強彈性與表達力。
Z-loss。 為了進一步提升訓練穩定性,我們在訓練目標中引入稱為 z-loss (de Brébisson and Vincent 2016) 的穩定項。z-loss 懲罰 logits 偏離零的程度,有效抑制模型產生可能導致數值不穩定或梯度爆炸的過大 logit 值。z-loss 定義為 logits 平方和:$\mathcal{L}_{\text{z-loss}}=\lambda\cdot\sum_{i}z_{i}^{2}$。我們發現 z-loss 對把梯度範數維持在健康範圍至關重要,在大量 GPU 節點上擴展訓練時尤其如此。經驗上,z-loss 係數 $\lambda=3\times 10^{-4}$ 能取得最佳平衡,有效穩定訓練而不損及模型表現。
5.2.2 規模化
本節描述讓自迴歸 WFM 高效規模化的技術。我們簡要分析模型的記憶體消耗、討論平行化策略,並與其他自迴歸模型的訓練配置做比較。
記憶體需求。 訓練期間,GPU 記憶體主要消耗於:
- 模型參數: 每參數 6 位元組。模型參數同時以 BF16 與 FP32 儲存。
- 梯度: 每參數 2 位元組。梯度以 BF16 儲存。
- 最佳化器狀態: 每參數 8 位元組。AdamW (Loshchilov and Hutter 2019) 的一階與二階動差皆以 FP32 儲存。
- 激活值: 約 $(2\times\text{number\_of\_layers}\times 17\times\text{seq\_len}\times\text{batch\_size}\times\text{d\_model})$ 位元組。關於最先進自迴歸模型激活記憶體的詳細分析,請參考 Korthikanti et al. 2023。
舉例來說,我們的 12B 模型(Cosmos-Predict1-12B)的參數、梯度與最佳化器狀態合計約需 192 GB 記憶體。由於超過單張 NVIDIA H100 GPU 的 80GB HBM3 容量,我們利用張量平行 (TP) (Shoeybi et al. 2019) 及其延伸序列平行 (SP) (Korthikanti et al. 2023),把記憶體需求與計算分散到多張 GPU。
張量平行 (TP)。 TP (Shoeybi et al. 2019) 把線性層的權重沿輸入或輸出特徵維度切分,切分方式以最小化 GPU 間通訊為準。例如在兩層前饋網路中,第一層權重沿輸出特徵維度分割,第二層權重沿輸入特徵維度分割。這種安排讓中間激活值能在本地處理,無需 GPU 間通訊;最終輸出再以 all-reduce 通訊合併。採用 TP 後,每張 GPU 只儲存線性層權重的 $1\mathbin{/}\text{TP\_SIZE}$。然而 TP 的預設實作對 LayerNorm 等操作仍沿序列維度複製激活值,造成冗餘。
序列平行 (SP)。 SP (Korthikanti et al. 2023) 延伸張量平行,進一步沿序列維度切分上下文。此方法適用於序列中各元素可獨立處理的運算子,例如自注意力層中的 LayerNorm 與 Dropout。啟用 SP 後,每張 GPU 只儲存 $1\mathbin{/}\text{TP\_SIZE}$ 的激活值。
與其他自迴歸模型的比較。 相較於流行的 LLM,我們的模型未採用 MQA 或 GQA 等節省記憶體的注意力變體。除此之外,我們的自迴歸模型刻意設計得與 LLM (Brown et al. 2020; Dubey et al. 2024; Jiang et al. 2023; Adler et al. 2024; Team 2024b; Yang et al. 2024a) 架構高度相似,因為這種一致性帶來彈性與可擴展性。利用更多平行化(如上下文平行與管線平行)進一步擴大模型規模與上下文長度的實驗留待未來工作。
5.2.3 訓練策略
我們分多個階段對自迴歸 WFM 進行預訓練。
- 階段 1: 第一階段以影片預測目標訓練模型。給定第一影格作為輸入條件,模型被訓練來預測未來的影片影格。此任務使用 17 影格的上下文長度,即模型以第一影格為輸入預測 16 個未來影格。
- 階段 1.1: 此階段同樣做影片預測,但上下文長度增加到 34 影格。我們在時間維度使用 YaRN 延伸來增加 RoPE 的上下文長度。
- 階段 2: 訓練的第 2 階段為模型引入文字條件化。文字嵌入透過新初始化的交叉注意力層引入。模型以 34 影格上下文訓練。為了提升文字生成影片的能力,模型按 5.1.3 節所述以圖像與影片聯合資料訓練。使用圖像批次時,由於圖像的上下文長度遠小於影片,我們使用更大的批次大小。
所有模型都以固定空間解析度 $640\times 1024$ 訓練。
冷卻階段。 預訓練之後,我們以高品質資料進行「冷卻 (Cooling-down)」階段,類似 LLM 的訓練實務 (Dubey et al. 2024)。在此階段,我們在高品質圖像-影片配對上訓練,同時把學習率線性衰減到 $0$。冷卻階段共 $30{,}000$ 次迭代。
| 配置 | 4B | 5B-Video2World | 12B | 13B-Video2World |
|---|---|---|---|---|
| 層數 | $16$ | $16$ | $40$ | $40$ |
| 模型維度 | $4{,}096$ | $4{,}096$ | $5{,}120$ | $5{,}120$ |
| 交叉注意力層 | ✗ | ✓ | ✗ | ✓ |
| 基礎學習率 | $1\times 10^{-3}$ | $3\times 10^{-4}$ | $1\times 10^{-3}$ | $5\times 10^{-4}$ |
| 權重衰減 | $0.01$ | |||
| 學習率預熱 | 線性排程,$5{,}000$ 次迭代 | |||
| 激活函數 | SwiGLU | |||
| FFN 隱藏維度 | $14{,}336$ | |||
| 注意力頭數 | $32$ | |||
| Key / Value 頭數 | $8$ | |||
| Token 數 | $12{,}800$ | |||
| 詞彙表大小 | $64{,}000$ | |||
| 位置嵌入 | 3D RoPE($\theta=500{,}000$)+ 3D APE |
表 14: Cosmos-Predict1 模型的配置細節。
我們訓練了兩組自迴歸 WFM。先建立兩個基礎模型:容量分別為 4B 與 12B。它們是純粹的下一個影片 token 預測器,不接受文字提示詞輸入。接著從各基礎模型衍生出 Video2World 版本:加入交叉注意力層,讓下一個影片 token 預測能利用文字提示詞輸入。
- Cosmos-Predict1-4B: 4B transformer 模型,做下一個影片 token 預測。以多階段訓練目標的階段 1 與階段 1.1 訓練。
- Cosmos-Predict1-5B-Video2World: 從 Cosmos-Predict1-4B 衍生的 5B transformer 模型,另以多階段訓練目標的階段 2 訓練。
- Cosmos-Predict1-12B: 12B transformer 模型,做下一個影片 token 預測。以多階段訓練目標的階段 1 與階段 1.1 訓練。
- Cosmos-Predict1-13B-Video2World: 從 Cosmos-Predict1-12B 衍生的 13B transformer 模型,另以多階段訓練目標的階段 2 訓練。
5.2.4 邁向即時生成的推論最佳化
我們的 Cosmos 自迴歸 WFM 與 LLM 在架構上相似,因此能借用成熟的 LLM 推論最佳化技術來解決序列式解碼的瓶頸。我們遵循 PyTorch (Paszke et al. 2019) 的 gpt-fast 【腳註:https://github.com/pytorch-labs/gpt-fast】實作,結合 key-value 快取、張量平行與 torch.compile。
推測式解碼。 為了進一步加速自迴歸 WFM,我們採用 Medusa 推測式解碼 (Speculative Decoding) 框架 (Cai et al. 2024)。不同於需要獨立草稿模型的常見推測式解碼方法 (Leviathan et al. 2023) 或加速有限的免訓練方法 (Teng et al. 2024),Medusa 在 transformer 主幹上擴充額外的解碼頭 (Decoding Head),平行預測多個後續 token,再以拒絕取樣 (Rejection Sampling) 驗證這些推測的 token。如此便緩解了一次只能處理一個 token 的瓶頸,加速推論。我們展示了 Medusa 技術在視覺自迴歸加速上的潛力,且不犧牲生成輸出的品質。
在我們的實作中,我們透過在架構中引入 Medusa 頭來微調預訓練的自迴歸 WFM。這些頭策略性地插在最後一層 transformer 隱藏狀態之後,所有主幹參數與最終的反嵌入層 (Unembedding Layer) 在不同頭之間共享。每個 Medusa 頭是帶 SiLU 激活與殘差連接的單層 FFN。我們進一步把多個 Medusa 頭的權重矩陣合併為統一的 FFN,最大化 token 預測時的平行度。注意我們沒有使用 Cai et al. 2024 的樹狀注意力機制。
為了探究自迴歸 WFM 的最佳 Medusa 配置,我們從兩方面深入研究:(1) 微調哪些 transformer 層,(2) 加多少個 Medusa 頭。針對第一個問題,我們比較全參數微調與選擇性凍結層。我們觀察到只微調 Medusa 頭會使多 token 預測表現不佳,而全參數微調會造成品質下降。我們的經驗顯示,解凍最後兩層 transformer 與最終反嵌入層、其餘主幹保持凍結,能取得最佳表現。此策略確保 Medusa 訓練達到不錯的推測式解碼準確率,又不會災難性遺忘 (Catastrophic Forgetting)。
| 模型 | Medusa 頭數 | 0 | 3 | 6 | 9 | 12 |
|---|---|---|---|---|---|---|
| 4B | Token 吞吐量(tokens/s) | 444.95 | 663.51 | 829.59 | 894.67 | 890.64 |
| 前向傳播次數 | 7680 | 2860 | 2073 | 1812 | 1682 | |
| 5B | Token 吞吐量(tokens/s) | 303.61 | 659.94 | 758.58 | 982.77 | 978.80 |
| 前向傳播次數 | 10240 | 2857 | 2382 | 1799 | 1673 |
表 15:Medusa 頭數對平均 token 吞吐量與前向傳播次數的影響。 實驗在 8 $\times$ H100 GPU 上進行,使用 50 部 $640\times 1024$ 解析度的未見測試影片。
為了探索最佳 Medusa 頭數,我們計算不同 Medusa 頭數下的模型 token 吞吐量與前向傳播次數。消融實驗在 8 $\times$ H100 GPU 上進行,以 50 部 $640\times 1024$ 解析度的未見測試影片評估。表 15 的結果顯示 Medusa 框架能有效加速推論:4B 模型最高達 $2.0\times$ token 吞吐量、前向傳播次數減少 $4.6\times$;5B 模型最高達 $3.2\times$ token 吞吐量、前向傳播次數減少 $6.1\times$。我們也顯示,雖然更多 Medusa 頭能減少生成所需的前向傳播次數,卻可能拖慢整體 token 吞吐量。我們發現 $9$ 個 Medusa 頭在計算效率與模型表現之間取得最佳權衡。
| 模型 | GPU 數 | 無 DD (s) | 無 DD+Medusa (s) | 有 DD (s) | 有 DD+Medusa (s) | VRAM (GB) |
|---|---|---|---|---|---|---|
| 4B | 1 | 31.04 | 23.52 | 61.49 | 53.08 | 29 |
| 4 | 18.20 | 13.87 | 29.60 | 25.63 | 31 | |
| 8 | 17.62 | 9.91 | 30.30 | 22.83 | 34 | |
| 5B | 1 | 39.68 | 24.97 | 70.39 | 54.72 | 59 |
| 4 | 25.59 | 20.96 | 37.29 | 33.35 | 51 | |
| 8 | 25.70 | 11.67 | 38.41 | 24.35 | 49 | |
| 12B | 1 | 84.78 | — | 116.66 | — | 45 |
| 4 | 47.49 | — | 60.27 | — | 36 | |
| 8 | 45.69 | — | 58.81 | — | 37 | |
| 13B | 1 | 109.18 | — | 140.24 | — | 77 |
| 4 | 67.80 | — | 80.76 | — | 55 | |
| 8 | 67.22 | — | 80.93 | — | 55 |
表 16: Cosmos 自迴歸模型在 $640\times 1024$ 解析度測試影片上的效能分析。【譯註:DD 指擴散解碼器 (Diffusion Decoder)。】
表 16 呈現整合 Medusa 後自迴歸 WFM 的效能分析。此分析在 H100 GPU 上進行,以 BF16 精度、$640\times 1024$ 解析度的測試影片評估。結果顯示 Medusa 實作在不同 GPU 配置下都能持續加速 4B 與 5B 模型的推論。
面向即時推論的低解析度適應。 我們透過把模型適應到較低的空間解析度 $320\times 512$ 來追求即時推論,這會降低每部影片的 token 數。具體來說,我們先用目標物理 AI 領域的影片,在 320p 低解析度影片上微調離散影片 tokenizer(第 4 節的 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p);接著以這個低解析度 tokenizer,在目標物理 AI 領域 $320\times 512$ 解析度的影片上微調原先以 $640\times 1024$ 解析度預訓練的自迴歸 WFM(5.2.3 節的 Cosmos-Predict1-4B);最後為微調後的低解析度自迴歸 WFM 加上 Medusa 頭。
| 模型($320\times 512$) | Token 吞吐量(tokens/s) | 影片吞吐量(frames/s) |
|---|---|---|
| Cosmos-Predict1-4B(含 Medusa) | 806.61 | 10.08 |
表 17: 低解析度適應後 Cosmos-Predict1-4B 的解碼吞吐量,在 8 $\times$ H100 80GB GPU 上以物理 AI 領域 $320\times 512$ 解析度、10 FPS 的影片測試。
我們在 8 $\times$ H100 GPU 上以 torch.compile 的「max-autotune」模式、BF16 精度進行推論基準測試,並以目標物理 AI 領域的 10-FPS 輸入影片評估。表 17 報告此配置下的平均 token 吞吐量與影格生成速度。我們觀察到模型能在 1 秒內生成 10 個影片影格,證明可以達成 10 FPS 的即時影片生成。
5.2.5 擴散解碼器
Cosmos tokenizer 使用輕量的編碼器-解碼器架構進行激進壓縮,減少 WFM 訓練的 token 數。激進壓縮的後果是有時會導致影片生成的模糊與可見瑕疵,在自迴歸 WFM 設定下尤其明顯——離散 token 化只用少量整數就要表徵內容豐富的影片。我們借助擴散解碼器設計 (Ramesh et al. 2022; OpenAI 2024a) 解決這項限制。具體來說,我們透過微調 5.1 節的 Cosmos-Predict1-7B-Text2World,打造一個更強大的 tokenizer 解碼器。

圖 15:Cosmos 擴散解碼器訓練。 訓練時,每部輸入影片會被 token 化兩次:一次經目標離散 tokenizer(DV8$\times$16$\times$16),另一次經約束較少的連續 tokenizer(CV8$\times$8$\times$8)。離散 token 影片作為擴散去噪器的條件輸入。

圖 16:Cosmos 擴散解碼器推論。 推論時,Cosmos-Predict1 模型輸出的影片 token 作為條件輸入提供給去噪器。
圖 15 說明我們如何為自迴歸 WFM 訓練擴散解碼器。對每部訓練影片,我們分別用 Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 與 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p 計算連續 token 影片與對應的離散 token 影片。我們注意到,得益於較溫和的連續 token 化過程與較不激進的壓縮方案($8\times 8\times 8$ 而非 $8\times 16\times 16$),Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 能產生比 Cosmos-Tokenize1-DV8$\times$16$\times$16-720p 更高品質的影片輸出。
離散 token 影片被當作 Cosmos-Predict1-7B 模型去噪器的條件輸入。為了計算條件輸入,我們先以可學習的詞彙嵌入層把離散 token 影片中的每個離散 token 嵌入為 16 維向量,接著沿 $x$ 與 $y$ 方向把嵌入上採樣 $2\times$,使條件輸入與來自連續 token 影片的雜訊輸入尺寸相同。我們把帶雜訊的連續輸入與條件輸入沿通道維度串接,作為擴散去噪器的輸入。去噪器的第一層做了通道維度擴充以容納新的輸入形狀。我們透過去除加入的雜訊來微調更新後的 Cosmos-Predict1-7B。由於離散 token 影片未受雜訊污染,去噪器學會利用條件輸入中蘊含的資訊來去噪。結果是一個為 tokenizer 服務的更高品質解碼器,它透過求解反向擴散問題來解碼離散 token。
圖 16 說明推論流程。自迴歸 WFM 輸出的離散 token 影片($8\times 16\times 16$ 離散壓縮)經兩步解碼為影片:第一步,運行條件式去噪器,根據自迴歸 WFM 的輸出生成連續 token 影片($8\times 8\times 8$ 連續壓縮);第二步,連續 token 影片由 Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 解碼,產生最終的 RGB 影片。
| 條件影格 $0$ | 影格 $10$ | 影格 $20$ | 影格 $30$ | |
|---|---|---|---|---|
| 4B | ![]() |
![]() |
![]() |
![]() |
| 12B | ![]() |
![]() |
![]() |
![]() |
提示詞:無。
| 條件影格 $0$ | 影格 $10$ | 影格 $20$ | 影格 $30$ | |
|---|---|---|---|---|
| 5B | ![]() |
![]() |
![]() |
![]() |
| 13B | ![]() |
![]() |
![]() |
![]() |
提示詞(節錄翻譯):一部汽車向前行駛、穿過大型高架橋下的影片。路面淨空,遠處可見幾輛其他車輛。天氣晴朗、時間為白天。場景是一條繁忙的高速公路,兩側有混凝土結構與綠意。
圖 17:Cosmos 自迴歸世界基礎模型生成的影片。 上方兩列為 4B 與 12B 模型的影片生成結果,下方兩列為帶文字提示詞的 Video2World 結果。我們觀察到,無論有無提示詞,12B 與 13B 模型都比 4B 與 5B 模型生成更銳利的影片與更好的運動。完整影片與更多範例請造訪我們的網站。
| 條件影格 $0$ | 影格 $10$ | 影格 $20$ | 影格 $30$ |
|---|---|---|---|
![]() |
![]() |
![]() |
![]() |
Cosmos-Predict1-13B-Video2World 的輸出
| 條件影格 $0$ | 影格 $10$ | 影格 $20$ | 影格 $30$ |
|---|---|---|---|
![]() |
![]() |
![]() |
![]() |
Cosmos-Predict1-13B-Video2World + 擴散解碼器的輸出
圖 18:擴散解碼器比較。 上方面板為 Cosmos-Predict1-13B-Video2World 模型的影片生成結果;下方面板為自迴歸模型輸出經擴散解碼器處理後的強化影片。我們觀察到,單獨的自迴歸模型產生模糊的結果,而擴散解碼器能在保留內容的同時強化影片的銳利度。
5.2.6 結果
圖 17 展示不同模型規模的自迴歸 WFM 的質性結果。在無提示詞設定下,比較 Cosmos-Predict1-4B 與 Cosmos-Predict1-12B 模型,我們觀察到 12B 模型生成的影片運動更佳、細節更銳利。同樣地,在有提示詞設定下,比較 Cosmos-Predict1-5B-Video2World 與 Cosmos-Predict1-13B-Video2World,可見 13B 模型的運動優於 5B 模型。
圖 18 展示使用擴散解碼器所獲得的增強效果。自迴歸模型的輸出偏模糊,主因是離散 tokenizer 的有損壓縮;使用擴散解碼器能在保留內容的同時強化細節。
我們從實驗中發現,5.1.5 節討論的提示詞上採樣器並未改善基於自迴歸的 Text2World WFM 的輸出。我們推測可能是因為這些 WFM 大部分訓練時間都在做純影片生成任務的預訓練,沒有被足夠強烈地要求利用文字輸入。
| 條件影格 $0$ | 影格 $10$ | 影格 $20$ | 影格 $30$ |
|---|---|---|---|
![]() |
![]() |
![]() |
![]() |
Cosmos-Predict1-4B 的輸出
圖 19:Cosmos 自迴歸 WFM 的失敗案例。 我們在生成影片中觀察到部分物體(以紅色標示)從下方意外出現的失敗案例。
| 模型 | 圖像條件化 | 影片條件化($9$ 影格) |
|---|---|---|
| Cosmos-Predict1-4B | $15\%$ | $1\%$ |
| Cosmos-Predict1-5B-Video2World | $7\%$ | $2\%$ |
| Cosmos-Predict1-12B | $2\%$ | $1\%$ |
| Cosmos-Predict1-13B-Video2World | $3\%$ | $0\%$ |
表 18: Cosmos 自迴歸模型的失敗率分析。
5.2.7 侷限
在自迴歸 WFM 的生成影片中,一個值得注意的失敗案例是物體從下方意外出現。圖 19 展示了此問題的一個例子。為了理解模型的失敗率,我們進行系統性研究:建立一個包含 $100$ 個物理 AI 輸入的評估集,對自迴歸 WFM 以兩種輸入模式——圖像(單影格)條件化與影片(9 影格)條件化——用所有模型生成影片。對所有生成影片,我們人工檢查失敗案例,並在表 18 報告失敗率。我們觀察到,較小的模型 Cosmos-Predict1-4B 與 Cosmos-Predict1-5B-Video2World 在單影格條件化下呈現較高的劣化率,而較大的模型 Cosmos-Predict1-12B 與 Cosmos-Predict1-13B-Video2World 更穩健。以 $9$ 影格影片條件化的生成對所有模型都穩定,失敗率低於 $2\%$。
5.3 評估
預訓練 WFM 是視覺世界模擬的通才,其能力應從多個面向衡量。在此我們從兩方面評估模型。第一,評估生成影片的 3D 一致性 (3D Consistency):理想的 WFM 應從幾何上合理的 3D 世界生成影片模擬。第二,評估生成影片的物理對齊 (Physics Alignment):計算呈現的動態遵循物理定律的程度。WFM 的評估絕非易事,我們也承認還有其他多個重要面向需要評估,更全面的評估留待未來工作。
5.3.1 3D 一致性
WFM 旨在透過影片生成模擬 3D 世界,因此評估生成影片與視覺世界 3D 結構的一致程度至關重要。除了看起來逼真,生成影片還應在時間中維持與場景物理原理的連貫性,這是下游物理 AI 應用的關鍵要求。
測試資料與基線模型。 為了能用基於多視角幾何 (Multi-view Geometry) 的既有工具有效衡量影片的 3D 一致性,我們聚焦於靜態場景的情境。我們從 RealEstate10K 資料集 (Zhou et al. 2018) 的測試集中隨機挑選 500 部影片作為資料集,並用專有 VLM 為影片生成描述,取得把影片描述為靜態場景的文字提示詞,如此計算指標時便無需考慮場景運動。我們以 VideoLDM (Blattmann et al. 2023b) 作為基線方法進行比較。
指標。 生成影片實質上是底層 3D 視覺世界的 2D 投影。我們設計下列指標來衡量生成影片的 3D 一致性:
- 幾何一致性。 我們透過量化對極幾何 (Epipolar Geometry) 約束的滿足程度來評估生成世界的 3D 一致性,包括 Sampson 誤差 (Sampson 1982; Hartley and Zisserman 2003) 與相機姿態估計演算法 (Schönberger and Frahm 2016; Schönberger et al. 2016) 在生成影片上的成功率。
- 視角合成一致性。 我們評估世界基礎模型在內插的新視角 (Novel Viewpoint) 合成圖像、同時與底層 3D 結構保持連貫的能力。
Sampson 誤差是一個興趣點到其在另一視角中對應對極線距離的一階近似。給定一對影格中的 $N$ 組點對應(以齊次座標表示)$\{\left(\bar{\mathbf{x}}_{i},\bar{\mathbf{y}}_{i}\right)\}_{i=1}^{N}$,Sampson 誤差定義為:
$$\epsilon_{\text{samp}}=\frac{1}{N}\sum_{i=1}^{N}\frac{\|\bar{\mathbf{y}}_{i}^{\top}\mathbf{F}\bar{\mathbf{x}}_{i}\|}{\sqrt{\left\|\mathbf{S}\mathbf{F}\bar{\mathbf{x}}_{i}\right\|_{2}^{2}+\left\|\mathbf{S}\mathbf{F}^{\top}\bar{\mathbf{y}}_{i}\right\|_{2}^{2}}}\;,\quad\text{where}\;\mathbf{S}=\begin{bmatrix}1&0&0\\ 0&1&0\\ 0&0&0\end{bmatrix},\tag{10}$$
其中 $\mathbf{F}$ 是由對應點估計的基礎矩陣 (Fundamental Matrix)。我們使用誤差函數的平方根版本,讓指標以像素為單位更直觀。我們結合 SuperPoint (DeTone et al. 2018) 與 LightGlue (Lindenberger et al. 2023) 從影格對偵測並匹配關鍵點對應,並以 OpenCV 的 8 點 RANSAC 演算法估計 $\mathbf{F}$。平均誤差以 $960\times 540$ 畫布的影格對角線長度正規化。
我們也以生成影片自我合成新視角的能力來評估其 3D 一致性。遵循新視角合成文獻的常見做法 (Mildenhall et al. 2020),我們每 8 個影格保留一個作為測試影格,用其餘訓練影格以 Nerfstudio 函式庫 (Tancik et al. 2023) 的預設設定擬合 3D 高斯潑濺 (3D Gaussian Splatting) 模型 (Kerbl et al. 2023)。我們報告 PSNR、SSIM 與 LPIPS (Zhang et al. 2018) 作為量化合成測試視角品質的指標。
| 幾何一致性 | 視角合成一致性 | ||||
|---|---|---|---|---|---|
| 方法 | Sampson 誤差 $\downarrow$ | 姿態估計成功率 (%) $\uparrow$ | PSNR $\uparrow$ | SSIM $\uparrow$ | LPIPS $\downarrow$ |
| VideoLDM (Blattmann et al. 2023b) | 0.841 | 4.4% | 26.23 | 0.783 | 0.135 |
| Cosmos-Predict1-7B-Text2World | 0.355 | 62.6% | 33.02 | 0.939 | 0.070 |
| Cosmos-Predict1-7B-Video2World | 0.473 | 68.4% | 30.66 | 0.929 | 0.085 |
| Cosmos-Predict1-4B | 0.433 | 35.6% | 32.56 | 0.933 | 0.090 |
| Cosmos-Predict1-5B-Video2World | 0.392 | 27.0% | 32.18 | 0.931 | 0.090 |
| 真實影片(參考) | 0.431 | 56.4% | 35.38 | 0.962 | 0.054 |
表 19: 基礎 Cosmos 模型的 3D 一致性評估。
結果。 量化評估結果見表 19。無論在幾何一致性或視角合成一致性方面,Cosmos WFM 都顯著優於基線模型。Cosmos WFM 的興趣點不僅更符合 3D 一致性,相機姿態估計成功率也明顯更高,反映整體品質與 3D 一致性的雙重提升,甚至達到真實世界影片的水準。在相機姿態成功估計的案例中,合成的保留視角在所有圖像合成指標上都展現更高品質。這些結果凸顯了 Cosmos WFM 生成 3D 一致影片的能力,確立其作為有效世界模擬器的地位。
5.3.2 物理對齊
理想的 WFM 應對物理定律有深刻理解,並產生遵循物理定律的未來觀察。雖然我們的預訓練 WFM 展現一定程度的物理理解並推進了最先進水準,但仍然很容易生成不遵守物理定律的例子。我們認為需要在資料整理中加入移除物理上不合理影片的額外步驟,並改進模型設計。我們把高度物理對齊的 WFM 留作未來工作,但仍有興趣衡量大規模資料驅動的預訓練能自然湧現多少直覺物理 (Intuitive Physics)。
為此,我們受 (Kang et al. 2024) 啟發,使用物理模擬引擎設計了一個受控的基準資料集。我們生成以物理為基礎的模擬,測試預訓練 WFM 對牛頓物理與剛體動力學 (Rigid Body Dynamics) 的遵循程度。具體來說,我們用模擬產生針對特定物理定律的測試情境的物理正確、照片級擬真影片。這些參考「真值 (Ground Truth)」影片再與 WFM 在共享上下文(過去觀察與擾動)下產生的「預測」影片比較。
合成資料生成。 我們使用 PhysX (NVIDIA 2024c) 與 Isaac Sim (NVIDIA 2024a) 設計八種 3D 情境,評估不同的物理效應:
- 自由落體: 物體落在平面上(重力、碰撞等)
- 傾斜平面斜坡: 物體沿斜面滾下(重力、轉動慣量等)
- U 形斜坡: 物體沿 U 形斜坡滾下(位能、動能等)
- 穩定堆疊: 處於平衡的物體堆(力的平衡)
- 不穩定堆疊: 失衡的物體堆(重力、碰撞等)
- 骨牌: 矩形磚塊依序倒下(動量傳遞、碰撞等)
- 蹺蹺板: 蹺蹺板兩側的物體(力矩、轉動慣性等)
- 陀螺儀: 平面上旋轉的陀螺(角動量、進動等)
對每種情境,我們隨機化動態物體的數量與類型(不同尺寸、紋理、形狀,選自 Omniverse 資產 (NVIDIA 2024b))以及背景外觀。我們模擬物體隨時間的運動狀態,並從 4 個不同的靜態相機視角渲染輸出影片。我們總共渲染了 800 部長 100 影格的 1080p 影片。每次模擬運行中的物體位置經過安排,使它們從第一影格起全部可見,以避免存在性的歧義。
傾斜平面斜坡——物體沿斜面滾下
| $t=0$(條件) | $t=11$ | $t=22$ | $t=32$ | |
|---|---|---|---|---|
| 模擬 | ![]() |
![]() |
![]() |
![]() |
| WFM | ![]() |
![]() |
![]() |
![]() |
U 形斜坡——兩物體從弧形斜坡兩端滾下
| $t=0$(條件) | $t=11$ | $t=22$ | $t=32$ | |
|---|---|---|---|---|
| 模擬 | ![]() |
![]() |
![]() |
![]() |
| WFM | ![]() |
![]() |
![]() |
![]() |
不穩定堆疊——物體堆因力失衡而倒塌
| $t=0$(條件) | $t=11$ | $t=22$ | $t=32$ | |
|---|---|---|---|---|
| 模擬 | ![]() |
![]() |
![]() |
![]() |
| WFM | ![]() |
![]() |
![]() |
![]() |
圖 20:模擬 vs 預訓練 WFM 的物理情境運行。 我們展示三個複雜度遞增的示例情境,各組第一列為參考(物理正確)模擬結果,第二列為 Cosmos-Predict1-7B-Video2World 的運行結果。WFM 以 9 個影格與聚焦於模擬物體運動狀態的提示詞為條件。每個例子顯示一個被追蹤的物體(藍色框與遮罩),用於計算物件層級指標(平均 IoU)。
指標。 我們的興趣在於:比較模擬真值影片與 WFM 直接生成的輸出,評估其對物理定律的遵循程度。因此,為了產生未來觀察,我們以真值影片的前幾個影格(1 或 9 個影格)作為 WFM 的條件。在適用時,我們額外以文字提示詞(用專有 VLM 對條件影格生成描述取得)作為 WFM 的條件,聚焦於過去觀察中被模擬物體的運動狀態。模擬與預測情境的一些範例見圖 20。評估使用下列指標:
- 像素層級指標。 像素層級的比較採用 PSNR 與 SSIM,比較 WFM 運行的預測影格與真值影片的參考影格。
- 特徵層級指標。 稍高層次的語意比較採用 DreamSim 相似度分數 (Fu et al. 2023)——一種特徵相似度指標——衡量預測與參考影格之間的相似性。
- 物件層級指標。 最後,由於我們最關心感興趣的物體如何受當前物理現象影響,我們利用追蹤來計算排除混淆因素(背景變化、視覺品質等)的物件層級指標。由於測試條件是合成生成的,我們擁有場景中動態物體的真值實例分割遮罩 (Instance Segmentation Mask)。我們使用 SAMURAI (Yang et al. 2024b) 把第一影格中的真值實例遮罩傳播到其餘預測影格以萃取軌跡,從而量化物件層級指標。我們對每一影格與每個感興趣物體計算真值與預測物體遮罩之間的交並比 (Intersection-over-Union, IoU)。
我們對影片中的影格、評估集中的影片,以及運行的四個隨機種子取平均。PSNR 與 SSIM 在所有影格上計算,但排除用作條件的影格。
| 像素層級 | 特徵層級 | 物件層級 | |||
|---|---|---|---|---|---|
| 模型 | 條件 | PSNR $\uparrow$ | SSIM $\uparrow$ | DreamSim $\uparrow$ | 平均 IoU $\uparrow$ |
| Cosmos-Predict1-7B-Video2World | 提示詞 + 1 影格 | 17.34 | 0.538 | 0.836 | 0.332 |
| Cosmos-Predict1-7B-Video2World | 提示詞 + 9 影格 | 21.06 | 0.691 | 0.859 | 0.592 |
| Cosmos-Predict1-14B-Video2World | 提示詞 + 1 影格 | 16.81 | 0.521 | 0.836 | 0.338 |
| Cosmos-Predict1-14B-Video2World | 提示詞 + 9 影格 | 20.21 | 0.635 | 0.860 | 0.598 |
| Cosmos-Predict1-4B | 1 影格 | 17.91 | 0.486 | 0.827 | 0.394 |
| Cosmos-Predict1-4B | 9 影格 | 18.13 | 0.482 | 0.859 | 0.481 |
| Cosmos-Predict1-5B-Video2World | 提示詞 + 1 影格 | 17.67 | 0.478 | 0.818 | 0.376 |
| Cosmos-Predict1-5B-Video2World | 提示詞 + 9 影格 | 18.29 | 0.481 | 0.864 | 0.481 |
| Cosmos-Predict1-12B | 1 影格 | 17.94 | 0.486 | 0.829 | 0.395 |
| Cosmos-Predict1-12B | 9 影格 | 18.22 | 0.487 | 0.869 | 0.487 |
| Cosmos-Predict1-13B-Video2World | 提示詞 + 1 影格 | 18.00 | 0.486 | 0.830 | 0.397 |
| Cosmos-Predict1-13B-Video2World | 提示詞 + 9 影格 | 18.26 | 0.482 | 0.865 | 0.482 |
表 20:物理對齊結果。 我們以像素層級、特徵層級與物件層級指標比較 Cosmos WFM 各變體對物理情境未來預測的準確度。指標在 33 個影格上計算,這是 Cosmos WFM 自迴歸變體支援的最大長度。
結果。 物理對齊的量化結果見表 20。根據量化與質性結果,我們有以下觀察。不意外地,以更多影格作為條件輸入時,模型能更準確預測整體物體運動學(因為能更好地推斷速度、加速度等一階與二階物理量)。
從表中我們也發現,在 9 影格條件設定下,擴散 WFM 在像素層級預測上優於自迴歸 WFM。這與我們的視覺觀察一致:基於擴散的 WFM 渲染的影片視覺品質更高。我們也注意到,結果並未顯示更大的模型在物理對齊上表現更好。雖然我們觀察到較大的模型渲染的影片視覺品質更高,但所有 WFM 在物理遵循上同樣吃力,需要更好的資料整理與模型設計。
更廣泛地說,我們觀察到上述剛體模擬已經在測試 WFM 的極限,是找出特定失敗案例的寶貴工具。失敗案例的範圍從低階問題(如物體恆存性喪失——物體自發出現與消失——以及形變)到更複雜的問題(如不合理的運動學、違反重力等)。我們相信這類結構化模擬提供了測試物理對齊的有用方法論。因此我們打算持續改進:納入更複雜的情境、提升照片級擬真度以彌合模擬到真實 (Sim-to-Real) 的差距(因為 WFM 預訓練資料是真實影片),並精進評估指標,以更全面地評估物理理解。
6 後訓練世界基礎模型
本節展示如何微調 Cosmos WFM 以支援多樣的物理 AI 應用。我們的範例包括:以相機控制後訓練 WFM,實現 3D 可導覽的視覺世界生成;在兩種不同的機器人設置上以動作控制後訓練 WFM,完成兩種不同的機器人操作任務;以及以多視角支援後訓練 WFM,用於訓練自動駕駛代理。
| 章節 | 模型 | 條件 |
|---|---|---|
| 6.1 節 | Cosmos-Predict1-7B-Video2World-Sample-CameraCond | 文字 + 圖像 + 相機 |
| 6.2 節 | Cosmos-Predict1-7B-Video2World-Sample-Instruction | 文字 + 影片 |
| 6.2 節 | Cosmos-Predict1-5B-Video2World-Sample-Instruction | 文字 + 影片 |
| 6.2 節 | Cosmos-Predict1-7B-Video2World-Sample-ActionCond | 動作 + 影片 |
| 6.2 節 | Cosmos-Predict1-5B-Video2World-Sample-ActionCond | 動作 + 影片 |
| 6.3 節 | Cosmos-Predict1-7B-Text2World-Sample-MultiView | 文字 |
| 6.3 節 | Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond | 文字 + 軌跡 |
| 6.3 節 | Cosmos-Predict1-7B-Video2World-Sample-MultiView | 文字 + 影片 |
表 21: 第 6 節討論的後訓練 WFM 一覽。
表 21 列出本節各小節討論的後訓練 WFM,並列出條件輸入以突顯其運作模式。注意每個模型名稱都加上「-Sample」,強調我們的目標是提供預訓練 WFM 的應用範例。這些模型絕非任何真實世界應用的完整系統或生產模型。開發者需要在自己的客製化資料集上,為其物理 AI 設置與目標應用微調這些 WFM。
6.1 相機控制的 WFM 後訓練
透過相機姿態條件化,我們把相機控制整合進 Cosmos-Predict1-7B-Video2World,使其成為有效的 3D 世界模擬器。我們把所得的後訓練 WFM 命名為 Cosmos-Predict1-7B-Video2World-Sample-CameraCond。我們聚焦於從單一參考輸入圖像生成 3D 世界,利用相機控制從指定的相機軌跡產生時間上連貫且 3D 一致的影片模擬,其中視角變化與場景的底層 3D 結構對齊。
6.1.1 資料集
此任務使用 DL3DV-10K (Ling et al. 2024)——一個大規模靜態場景影片資料集。前處理步驟中,我們把所有影片切成 256 影格的片段。為了對片段中所有影格取得稠密的相機姿態標註,我們用 GLOMAP (Pan et al. 2025) 對切好的片段執行運動恢復結構 (Structure-from-Motion)。我們把第一影格的相機姿態設為恆等變換,並計算所有後續影格的相對相機姿態。我們也用專有 VLM 為影片生成描述,取得把影片描述為靜態場景的文字提示詞。
6.1.2 微調
我們把取樣的潛在嵌入與 Plücker 嵌入 (Sitzmann et al. 2021) 串接來加入相機控制條件化,Plücker 嵌入與潛在嵌入具有相同的空間維度。具體來說,給定相機姿態,我們透過下式計算 Plücker 座標:
$$\mathbf{r}=(\mathbf{d},\mathbf{m})\in\mathbb{R}^{6}\;\;\;\text{where}\;\;\mathbf{m}=\mathbf{c}\times\mathbf{d}\;,\tag{11}$$
其中 $\mathbf{c}$ 是相機中心位置,$\mathbf{d}$ 是每個潛在像素的單位光線方向(潛在嵌入被視為下採樣的圖像)。所有相機姿態都是相對於初始影格的相對姿態。Cosmos-Predict1-7B-Video2World 使用的 Cosmos-Tokenize1-CV8$\times$8$\times$8-720p 具有 $8\times$ 的時間壓縮率,因此每 8 個影格,我們取第 4 影格的 Plücker 嵌入與對應的潛在表徵串接。
我們把訓練影片的輸入影格縮放為 $704\times 1252$,並以鏡射填補至 $704\times 1280$。訓練時取樣 57 個影格。訓練目標與其他超參數與基礎擴散 WFM 訓練(5.1.3 節)相同。
6.1.3 評估
我們假設給定世界的單一參考圖像,從輸入圖像生成未來運行的影片。我們與 CamCo (Xu et al. 2024)——此設定下相機可控影片生成的最先進模型——比較。為求公平,我們使用同樣在 DL3DV-10K (Ling et al. 2024) 訓練集上微調過的 CamCo 模型。由於我們的後訓練 WFM 生成 57 影格而 CamCo 只能生成 14 影格,我們比較相同的 57 影格軌跡,並對 CamCo 做 $4\times$ 時間下採樣。CamCo 的影片解析度上限為 $256\times 256$,我們另外對輸入圖像與測試影格做最大置中裁切後評估。
測試資料使用 5.3.1 節描述的 RealEstate10K (Zhou et al. 2018) 測試集中的相同 500 個樣本。我們以初始影格為參考圖像、以資料集提供的相機軌跡為相機控制輸入,並額外重新縮放,使軌跡兩端點之間的距離正規化為 1。
指標。 遵循 Xu et al. 2024,我們從兩方面評估後訓練世界模型的相機可控性:影片生成品質與 3D 一致性。影片品質方面,我們用 FID (Heusel et al. 2017) 與 FVD (Unterthiner et al. 2019) 分別評估影格層級與影片層級的品質。我們用相同的測試資料作為參考影片來計算指標(注意並非用於像素層級比較)。
3D 一致性方面,我們透過運動恢復結構函式庫 (Schönberger and Frahm 2016; Schönberger et al. 2016; Pan et al. 2025) 重新估計相機姿態的能力來評估,並把結果與輸入的相機控制軌跡比較。給定影片的 $N$ 個影格,我們把相機軌跡誤差量化為兩項:平均旋轉誤差 $\epsilon_{\text{rot}}$ 與平移誤差 $\epsilon_{\text{trans}}$,分別定義為:
$$\epsilon_{\text{rot}}=\frac{1}{N}\sum_{i=1}^{N}\cos^{-1}\!\left(\frac{\mathrm{trace}(\mathbf{\hat{R}}_{i}^{\top}\mathbf{R}_{i})-1}{2}\right)\qquad\text{and}\qquad\epsilon_{\text{trans}}=\frac{1}{N}\sum_{i=1}^{N}\left\|\mathbf{\hat{t}}_{i}-\mathbf{t}_{i}\right\|_{2}\;,\tag{12}$$
其中 $\mathbf{R}_{i}$ 與 $\mathbf{t}_{i}$ 是第 $i$ 影格的輸入旋轉與平移(作為真值),$\mathbf{\hat{R}}_{i}$ 與 $\mathbf{\hat{t}}_{i}$ 是重新估計的量。為了處理相機姿態估計結果存在相似變換 (Similarity Transformation) 不確定性的問題,我們遵循 Lin et al. 2021,對預測的相機軌跡執行 Procrustes 分析以對齊真值。

圖 21:相機控制模型的質性比較。 給定輸入影格與相機軌跡(顏色隨時間從紅到紫編碼),我們就生成的未來影格與重新估計的相機姿態,比較 Cosmos-Predict1-7B-Video2World-Sample-CameraCond 與 CamCo (Xu et al. 2024)。CamCo 受資料分布偏移之苦,常生成不準確的軌跡,甚至生成分布外 (Out-of-distribution) 的圖像、導致相機姿態無法估計。相較之下,Cosmos 相機控制模型能成功生成與相機控制輸入對齊的未來影格,同時維持高影片品質與 3D 一致性。
| 相機軌跡對齊 | 影片生成品質 | ||||
|---|---|---|---|---|---|
| 方法 | 姿態估計成功率 (%) $\uparrow$ | 旋轉誤差 ($\degree$) $\downarrow$ | 平移誤差 $\downarrow$ | FID $\downarrow$ | FVD $\downarrow$ |
| CamCo (Xu et al. 2024) | 43.0% | 8.277 | 0.185 | 57.49 | 433.24 |
| Cosmos-Predict1-7B-Video2World-Sample-CameraCond | 82.0% | 1.646 | 0.038 | 14.30 | 120.49 |
表 22: 相機控制後訓練 WFM 的量化比較。
比較。 結果見表 22。首先,我們的後訓練 WFM 能生成逼真且連貫的 3D 世界,這體現在更低的 FID/FVD 分數(更高的視覺品質)與更高的相機姿態估計成功率。Cosmos-Predict1-7B-Video2World-Sample-CameraCond 展現更好的相機控制,重新估計的相機軌跡顯著更貼近原始控制輸入。
我們也在圖 21 提供視覺比較。CamCo 難以生成輸入圖像以外的內容,而 Cosmos-Predict1-7B-Video2World-Sample-CameraCond 能有效生成符合 3D 世界結構的視覺內容。注意兩個模型都在 DL3DV-10K 上後訓練、在 RealEstate10K 資料集上評估,訓練與測試之間存在顯著的分布偏移。Cosmos 模型成功克服了這種分布偏移,也展現對未見輸入相機軌跡的泛化能力。
場景一(左為輸入影格,其餘為各方向生成結果,取影格 14、28、42、57):
| 控制 | 輸入影格 | 生成影格 |
|---|---|---|
| 向前移動 | ![]() |
![]() |
| 向後移動 | ![]() |
![]() |
| 向左旋轉 | ![]() |
![]() |
| 向右旋轉 | ![]() |
![]() |
場景二:
| 控制 | 輸入影格 | 生成影格 |
|---|---|---|
| 向前移動 | ![]() |
![]() |
| 向後移動 | ![]() |
![]() |
| 向左旋轉 | ![]() |
![]() |
| 向右旋轉 | ![]() |
![]() |
場景三:
| 控制 | 輸入影格 | 生成影格 |
|---|---|---|
| 向前移動 | ![]() |
![]() |
| 向後移動 | ![]() |
![]() |
| 向左旋轉 | ![]() |
![]() |
| 向右旋轉 | ![]() |
![]() |
圖 22:Cosmos-Predict1-7B-Video2World-Sample-CameraCond 的搖桿控制結果。 對每個輸入影格(最左欄),我們套用 4 種以類搖桿控制產生的相機軌跡:向前移動、向後移動、向左旋轉、向右旋轉。我們視覺化生成影片的第 14、28、42、57 影格。
不同隨機種子的生成結果(每組使用相同輸入影格與相機條件,各列為 3 個不同種子的生成影片,取影格 19、38、57):
| 場景 | 輸入影格 | 不同種子的生成影格(向後移動) |
|---|---|---|
| 1 | ![]() |
![]() |
| 2 | ![]() |
![]() |
| 3 | ![]() |
![]() |
| 場景 | 輸入影格 | 不同種子的生成影格(向右旋轉) |
|---|---|---|
| 1 | ![]() |
![]() |
| 2 | ![]() |
![]() |
| 3 | ![]() |
![]() |
圖 23:Cosmos-Predict1-7B-Video2World-Sample-CameraCond 在不同種子下的結果。 我們展示相機控制模型在相同輸入圖像與相機條件下模擬多樣未來的能力。每組套用相同的輸入影格與搖桿建立的相機條件:第一組為向後移動,第二組為向右旋轉。每組各欄展示 3 個不同隨機種子的生成影片。我們視覺化生成影片的第 19、38、57 影格。
質性結果。 圖 22 展示對相機做類搖桿控制輸入的結果,包括向前移動、向後移動、向左旋轉與向右旋轉。這展示了使用者能以搖桿控制模型生成未來影片影格、在模擬世界中導覽的應用情境。物理 AI 代理人也能利用此類控制來預測不同情境下世界的未來。
為了展示生成的多樣性,圖 23 呈現相同輸入圖像與相機控制、不同隨機種子下的生成結果。Cosmos-Predict1-7B-Video2World-Sample-CameraCond 能生成不同的世界,同時維持影片中的 3D 空間與時間連貫性。這可用於在給定當前狀態下模擬多種可能的未來。
6.2 機器人操作的 WFM 後訓練
世界模型有潛力成為機器人操作的強大規劃器與模擬器。此處我們展示如何為兩項任務微調預訓練 WFM:(1) 基於指令的影片預測,(2) 基於動作的下一影格生成。基於指令的影片預測中,輸入是機器人當前的影片影格與一段文字指令,輸出是機器人遵循該指令的預測影片。基於動作的下一影格預測中,輸入是機器人當前的影片影格與當前到下一影格之間的動作向量,輸出是機器人執行指定動作後的下一影格預測。給定一連串動作,模型可以自迴歸地運行,預測機器人執行這些動作的影片。
6.2.1 資料集
我們為上述兩項任務整理了兩個資料集。針對基於指令的影片預測,我們建立了名為 Cosmos-1X 的內部資料集,包含約 200 小時由 1x.Tech (Technologies 2024) 的人形機器人 EVE 拍攝的第一人稱影片,任務多樣,包括導航、摺衣服、清理桌面、撿拾物品等。我們從原始影片中挑選約 $12{,}000$ 段長度 1 到 9 秒的片段 (Episode)。每段都標註一句話的指令,之後再以專有 VLM 上採樣。影片以 30 FPS、$512\times 512$ 解析度拍攝。
針對基於動作的下一影格生成,我們使用公開的 Bridge 資料集 (Ebert et al. 2022),並採用與先前工作 (Zhu et al. 2024) 相同的配置以便比較。Bridge 資料集包含約 $20{,}000$ 段機械手臂在廚房環境中執行不同任務的第三人稱影片,解析度 $320\times 256$、5 FPS。對每個影片影格,對應動作定義為夾爪座標空間中的 7 維向量 $(\Delta x,\Delta y,\Delta z,\Delta\theta_{r},\Delta\theta_{p},\Delta\theta_{y},\Delta\mbox{Gripper})$,與 OpenVLA (Kim et al. 2024) 相同。
6.2.2 微調
我們對 Cosmos-Predict1-7B-Video2World(5.1 節)與 Cosmos-Predict1-5B-Video2World(5.2 節)都進行了基於指令的影片預測與基於動作的下一影格預測任務的微調。
針對基於指令的影片預測,我們基於基礎 WFM 建立兩個模型:第一個叫 Cosmos-Predict1-7B-Video2World-Sample-Instruction,第二個叫 Cosmos-Predict1-5B-Video2World-Sample-Instruction。我們計算指令的 T5 嵌入,並在基礎模型微調時經交叉注意力引入。
針對基於動作的下一影格預測,我們同樣基於基礎 WFM 建立兩個模型:第一個叫 Cosmos-Predict1-7B-Video2World-Sample-ActionCond,第二個叫 Cosmos-Predict1-5B-Video2World-Sample-ActionCond。
由於動作是預訓練時未曾遇過的新模態,我們在模型中引入額外的條件化模組。對 Cosmos-Predict1-5B-Video2World-Sample-ActionCond,我們加入一個動作嵌入器 (Action Embedder) MLP,把動作向量投影為張量,再經交叉注意力併入模型。對 Cosmos-Predict1-7B-Video2World-Sample-ActionCond,我們同樣加入動作嵌入器 MLP 把動作投影為張量,但改為把它加到 DiT 模組的時間戳嵌入 (Timestamp Embedding) 上來併入模型。
6.2.3 評估
圖 24:Cosmos-1X 資料集上基於指令的影片預測人工評估結果。(a)Cosmos-Predict1-7B-Video2World-Sample-Instruction vs VideoLDM-Instruction;(b)Cosmos-Predict1-5B-Video2World-Sample-Instruction vs VideoLDM-Instruction。結果顯示,相較基線模型(VideoLDM-Instruction),我們微調後的模型在四個評估維度上都獲得更高的偏好度。【譯註:原圖為長條圖,HTML 版未提供可下載之圖檔。】
Cosmos-Predict1-7B-Video2World-Sample-Instruction 的結果:
| 輸入影格 | 指令條件式生成 |
|---|---|
![]() |
![]() |
提示詞:Organize books by placing them vertically on a shelf.(把書直立放上書架整理好。)
| 輸入影格 | 指令條件式生成 |
|---|---|
![]() |
![]() |
提示詞:Fold a green fabric item on a table.(在桌上摺疊一件綠色織物。)
Cosmos-Predict1-5B-Video2World-Sample-Instruction 的結果:
| 輸入影格 | 指令條件式生成 |
|---|---|
![]() |
![]() |
提示詞:Grip and elevate a green object from a box on a tidy worktable.(從整潔工作桌上的箱子中抓取並舉起一個綠色物體。)
| 輸入影格 | 指令條件式生成 |
|---|---|
![]() |
![]() |
提示詞:Retrieve a box from a storage shelf using its articulated hands in a warehouse setting.(在倉儲環境中,用其關節式雙手從貨架上取回一個箱子。)
圖 25:Cosmos-1X 資料集上基於指令的影片預測範例。 左側為 Cosmos-Predict1-7B-Video2World-Sample-Instruction 模型的結果,右側為 Cosmos-Predict1-5B-Video2World-Sample-Instruction 模型的結果。
Cosmos-Predict1-7B-Video2World-Sample-ActionCond(左)與 Cosmos-Predict1-5B-Video2World-Sample-ActionCond(右):
| 輸入影格 | 預測影格 | |
|---|---|---|
| 預測 | ![]() |
![]() |
| 真值 | ![]() |
![]() |
| 輸入影格 | 預測影格 | |
|---|---|---|
| 預測 | ![]() |
![]() |
| 真值 | ![]() |
![]() |
圖 26:Bridge 資料集上基於動作的下一影格預測範例。 左側為 Cosmos-Predict1-7B-Video2World-Sample-ActionCond 模型的結果,右側為 Cosmos-Predict1-5B-Video2World-Sample-ActionCond 模型的結果。如圖所示,兩個模型的預測影片影格都與真值影片影格高度吻合。
針對基於指令的影片預測,我們在 Cosmos-1X 資料集上微調 VideoLDM (Blattmann et al. 2023b),得到 VideoLDM-Instruction 作為比較基線。為了評估模型的影片生成表現,我們定義以下維度:
- 指令遵循: 生成影片是否與輸入語言指令對齊?
- 物體恆存性: 場景中出現的物體是否貫穿整部生成影片持續存在?
- 真實性: 生成影片是否忠實呈現真實世界、沒有意外出現的想像物體?
- 整體: 生成影片是否合理到足以讓機器人據此規劃?
人工評估者觀看由不同模型以相同語言指令生成的一對匿名影片,並沿上述維度比較。十位人工評估者對 23 個測試片段進行評估,統計結果匯總於圖 24。
如圖所示,Cosmos-Predict1-7B-Video2World-Sample-Instruction 與 Cosmos-Predict1-5B-Video2World-Sample-Instruction 在四個評估維度上都優於 VideoLDM-Instruction。Cosmos-Predict1-7B-Video2World-Sample-Instruction 獲得 $78.3\%$ 的整體偏好,VideoLDM-Instruction 僅 $13.0\%$。Cosmos-Predict1-5B-Video2World-Sample-Instruction 也優於基於擴散的 VideoLDM-Instruction。兩個微調 WFM 的部分預測影片影格見圖 25,可見預測影片的品質。
針對基於動作的下一影格預測,我們在 Bridge 資料集上微調模型。作為基線,我們微調 IRASim (Zhu et al. 2024) 得到基於動作的下一影格預測模型 IRASim-Action。我們以自迴歸方式執行下一影格預測來生成影片。為了評估影片生成品質,我們在官方 Bridge 測試集中隨機挑選 100 個片段,把生成影片與真值影片比較。
| 方法 | PSNR $\uparrow$ | SSIM $\uparrow$ | Latent L2 $\downarrow$ | FVD $\downarrow$ |
|---|---|---|---|---|
| IRASim-Action | 19.13 | 0.64 | 0.38 | 593 |
| Cosmos-Predict1-5B-Video2World-Sample-ActionCond | 19.95 | 0.80 | 0.36 | 434 |
| Cosmos-Predict1-7B-Video2World-Sample-ActionCond | 21.14 | 0.82 | 0.32 | 190 |
表 23: Bridge 資料集上基於動作的下一影格預測評估。
計算所得的指標匯總於表 23,包括 PSNR、SSIM、Latent L2 (Zhu et al. 2024) 與 FVD。如表所示,Cosmos-Predict1-5B-Video2World-Sample-ActionCond 與 Cosmos-Predict1-7B-Video2World-Sample-ActionCond 模型都優於基線模型(IRASim-Action)。部分預測影片影格見圖 26,可見預測影片相對於真值的品質。
6.3 自動駕駛的 WFM 後訓練
真實駕駛場景的世界模型有潛力成為訓練自動駕駛代理的強大模擬引擎。由於大多數自動駕駛車輛配備多台朝向不同方向的相機,理想的自動駕駛世界模型也應是多視角 (Multi-view) 的,最好能精確匹配目標車輛的感測器配置。此處我們展示如何微調預訓練 WFM,建立面向自動駕駛任務的多視角世界模型。
6.3.1 資料集
我們整理了名為真實駕駛場景 (Real Driving Scene, RDS) 的內部資料集,包含約 360 萬段 20 秒的環景影片片段(約合 $20{,}000$ 小時的資料),由 NVIDIA 內部駕駛平台拍攝。每段片段錄自六個相機視角:前、左、右、後、左後、右後。此外,資料集包含自車運動 (Ego-motion) 資訊,我們用它建構軌跡資料。我們以前置相機影片的錄影時間戳同步其他所有視角的影格。
此資料集是從大型標註資料庫中挑選、以匹配目標資料屬性分布。具體的屬性標籤包括:
- 周遭車輛密度(如:無、低、中、高)
- 天氣(如:晴朗、下雨、下雪、起霧)
- 光照(如:白天、夜晚)
- 自車速度(如:靜止、低速、市區速度、高速公路速度)
- 自車行為(如:高、中、低曲率的軌跡與加速度)
- 道路類型/人口密度(依 OpenStreetMap 定義:鄉村、住宅區、都市)
此外,資料集透過第二輪資料探勘擴充,確保包含稀有道路結構(如收費站、橋梁、隧道、減速丘等)的片段達到最低數量。最後,各相機視角的影片分別加上描述,開頭是模板文字:「The video is captured from a camera mounted on a car. The camera is facing forward|left|right|backward|rear-left|rear-right.」
6.3.2 微調
我們使用 RDS 資料集,把 Cosmos-Predict1-7B-Text2World(5.1 節)微調為多視角世界模型。為了確保多視角間影片生成的一致性,我們對 5.1 節描述的架構設計做了些微修改,並微調 WFM 使其同時生成全部六台相機的影片。
我們建立了三個多視角世界模型,匯總於表 21。第一個叫 Cosmos-Predict1-7B-Text2World-Sample-MultiView,是能根據文字提示詞輸入生成六個相機視角的多視角世界模型。第二個叫 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond,建立在 Cosmos-Predict1-7B-Text2World-Sample-MultiView 之上,額外接受軌跡輸入作為條件訊號。最後一個模型 Cosmos-Predict1-7B-Video2World-Sample-MultiView 從多視角 Text2World 模型微調而來,以支援影片式條件化——把先前影格納入生成過程。Cosmos-Predict1-7B-Video2World-Sample-MultiView 能接收 Cosmos-Predict1-7B-Text2World-Sample-MultiView 的影片輸出並生成其延續。三個模型都輸出 6 個視角、57 影格、$848\times 480$ 解析度的影片。
視角獨立的位置嵌入與視角嵌入。 我們沒有把 FPS 感知 3D RoPE 位置嵌入擴充出額外的視角維度,而是選擇對每個視角獨立套用 5.1 節描述的相同位置嵌入。為了表達視角差異,我們修改去噪函數 $D_{\theta}$,使其接受額外的視角嵌入 (View Embedding) 作為輸入。也就是說,相機視角資訊由全域視角嵌入提供,而非位置嵌入。
視角相依的交叉注意力。 在多視角設定中,同一場景的六個視角各有不同的影片描述。雖然我們把六個視角整體視為擴散過程的狀態、在六個視角的所有元素之間做自注意力去噪,但我們發現對文字輸入採用視角相依的交叉注意力更有利。具體來說,每個視角的交叉注意力只關注該視角的文字描述。注意在我們的資料集中,每個視角都有不同的影片描述。透過視角嵌入與視角相依交叉注意力,我們微調 Cosmos-Predict1-7B-Text2World 得到 Cosmos-Predict1-7B-Text2World-Sample-MultiView。
軌跡控制條件。 作為選項,除了文字條件,我們微調模型使其產生符合給定未來軌跡路徑的影片,以實現對代理更精確的控制。這能生成既遵循真實世界資料記錄的駕駛軌跡、又符合輸入文字描述所指定駕駛環境的獨特駕駛情境。微調後的模型是 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond。
我們把軌跡定義為 3D 空間中的 64 個點序列,代表代理從初始位置 $(0,0,0)$ 到最終目的地的平移序列,每點間隔 0.1 秒。我們計算軌跡輸入的嵌入,把結果作為微調 Cosmos-Predict1-7B-Video2World 模型去噪器的條件輸入。我們注意到,遵循先前工作 (Kim et al. 2020; Kim et al. 2021; Hu et al. 2023) 或機器人操作任務(6.2 節)的做法,透過提供逐區間的動作向量可以達到更細緻的控制訊號。我們把這類延伸留給未來工作。
| 範例一 | 範例二 |
|---|---|
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
左側提示詞(節錄翻譯):影片捕捉高速公路場景,前景有一輛白色卡車朝鏡頭駛來。卡車有大型貨艙,後方跟著一位戴全罩式安全帽的機車騎士。道路標有白線,右側有金屬護欄。天空多雲時晴,路旁可見綠樹與灌木。從動態模糊與卡車、機車騎士的視角變化可知影片是從行進中的車輛拍攝。 右側提示詞(節錄翻譯):畫面顯示霧中高速公路上的多車連環事故。濃霧嚴重影響能見度,只能看見前方車輛的尾燈。突然間煞車燈亮起,車輛開始急轉並緊急停下。公路上滿是停下與撞毀的車輛,四周被霧籠罩,更添場面的混亂。
圖 27:由 Cosmos-Predict1-7B-Text2World-Sample-MultiView 生成、再由 Cosmos-Predict1-7B-Video2World-Sample-MultiView 延長至 8 秒的文字條件式範例。 本圖把六個相機視角以群組視覺化,每列對應一個特定時間點。左例描繪機車與大卡車並行的高速公路場景;右例顯示自車在大雪天跟隨一輛轎車右轉。
| 範例一 | 範例二 |
|---|---|
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
左側提示詞(節錄翻譯):高聳、設計精緻、由內部照亮的冰之城堡。前方天空展現絢麗的日落,明亮的月亮位於城堡左側,為場景灑上藍色調。快速移動、深沉而戲劇性的雲層增添超凡脫俗的氛圍。3D 寫實藝術風格著重光影與質感,營造出震撼的視覺效果。 右側提示詞(節錄翻譯):畫面捕捉海洋與嶙峋岩岸交會的海岸地帶。前方湛藍的海浪拍打岩石、激起白色泡沫。峭壁呈棕綠交錯的色調,顯示有植被,可能還有苔蘚或藻類。
圖 28:由 Cosmos-Predict1-7B-Text2World-Sample-MultiView 生成、再由 Cosmos-Predict1-7B-Video2World-Sample-MultiView 延長至 8 秒的文字條件式範例。 後訓練的世界模型有效保留了泛化能力。左例中自車駛向一座冰之城堡,右例中自車行駛在河面上。
6.3.3 評估
| 生成品質 | 多視角一致性 | |||
|---|---|---|---|---|
| 方法 | FID $\downarrow$ | FVD $\downarrow$ | TSE $\downarrow$ | CSE $\downarrow$ |
| VideoLDM-MultiView | 60.84 | 884.46 | 1.24 | 6.48 |
| Cosmos-Predict1-7B-Text2World-Sample-MultiView | 32.16 | 210.23 | 0.68 | 2.11 |
| Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond | - | - | 0.59 | 2.02 |
| 真實影片(參考) | - | - | 0.69 | 1.71 |
表 24: 後訓練多視角世界模型的多視角駕駛影片生成評估。
| 方法 | TAE-ATE $\downarrow$ | TAE-RPE-R $\downarrow$ | TAE-RPE-t $\downarrow$ | TFE $\downarrow$ |
|---|---|---|---|---|
| VideoLDM-MultiView | 0.88 | 22.94 | 0.77 | - |
| Cosmos-Predict1-7B-Text2World-Sample-MultiView | 0.77 | 4.25 | 0.29 | - |
| Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond | 0.54 | 4.31 | 0.18 | 20.20 |
| 真實影片(參考) | 0.49 | 4.60 | 0.14 | 13.49 |
表 25:後訓練多視角世界模型的多視角駕駛影片生成軌跡一致性評估。 TAE 數值為方便起見乘以 $10^{2}$,TFE 的單位為公分。
我們先在圖 27 呈現文字條件式的質性結果。使用 Cosmos-Predict1-7B-Text2World-Sample-MultiView 生成六視角的 57 影格影片,再以 Cosmos-Predict1-7B-Video2World-Sample-MultiView 模型延長至 201 影格。圖 28 展示預訓練世界模型如何增強泛化能力,得以生成 RDS 資料集中稀有或領域外的場景,例如在河面上行駛。最後,圖 29 展示 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 的結果,自車精確地遵循輸入軌跡。
在量化結果方面,作為基線,我們以相同的微調配方微調 VideoLDM (Blattmann et al. 2023b),得到名為 VideoLDM-MultiView 的多視角世界模型。我們使用一組評估指標,衡量影片生成品質、多視角一致性與軌跡遵循準確度。影片生成品質以 1000 個樣本計算分數。針對一致性相關指標,為了更好地理解各模型在不同情境下的行為,我們把真值軌跡分為四類:直行、左轉、右轉與其他(包含靜止或複雜運動)。每一類蒐集 200 個樣本及其對應的提示詞與條件,共 800 個樣本。以下詳述各指標與結果。
| 軌跡輸入視覺化 | 影格 25 | 影格 50 | 影格 75 | 影格 100 |
|---|---|---|---|---|
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
![]() |
圖 29:Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 的軌跡條件式生成範例。 給定最左欄的軌跡輸入,我們生成遵循該軌跡的多視角影片。本圖視覺化前置相機視角。
生成品質。 我們用 FID (Heusel et al. 2017) 與 FVD (Unterthiner et al. 2019) 衡量生成影片相對真實影片的品質。我們先對每個視角從每部影片抽取 16 個影格計算分數,再報告每個方法在所有視角的平均分數。如表 24 所示,Cosmos-Predict1-7B-Text2World-Sample-MultiView 與 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 在兩項指標上都顯著優於 VideoLDM-MultiView,證明我們預訓練的 7B 擴散 WFM 品質優於 VideoLDM-MultiView 基線。
多視角一致性。 我們用 5.3.1 節所述 Sampson 誤差 (Sampson 1982; Hartley and Zisserman 2003) 的延伸版本來量化生成多視角影片的幾何一致性。由於 RDS 資料集的真值影片共享相似的魚眼相機內參,我們以中位數校準把關鍵點去畸變 (Undistort) 到統一尺寸 $960\times 540$、水平視野 120 度的標準針孔相機。此設定下對生成多視角影片計算兩項指標:
- 時間 Sampson 誤差 (Temporal Sampson Error, TSE) 衡量每台相機生成的內容是否隨時間保持一致,取各視角相鄰影格 Sampson 誤差的中位數。
- 跨視角 Sampson 誤差 (Cross-view Sampson Error, CSE) 衡量多視角一致性是否隨時間維持,取不同生成視角間的 Sampson 誤差對時間平均。CSE 使用的基礎矩陣由所有時間影格累積的關鍵點估計。
如表 24 所示,Cosmos-Predict1-7B-Text2World-Sample-MultiView 與 Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 都比 VideoLDM-MultiView 呈現更佳的多視角幾何一致性。從我們的 WFM 微調而來的世界模型,生成影片的整體幾何合理性明顯更好。我們也注意到,加入軌跡控制條件後,得益於顯式的 3D 引導,一致性進一步提升——Cosmos-Predict1-7B-Text2World-Sample-MultiView-TrajectoryCond 排名最佳。
軌跡一致性:軌跡吻合誤差 (Trajectory Agreement Error, TAE)。 我們基於 Teed and Deng 2021 的形式化,設計了與 Liang et al. 2024 類似的穩健多視角相機姿態估計流程。此姿態估計流程具備線上動態遮罩生成模組與高效的稠密光束法平差 (Dense Bundle Adjustment) 模組,能穩健且即時地估計多視角相機姿態。我們用此流程分別以「前 + 左前」與「前 + 右前」兩種多視角相機組合估計前置相機的姿態,再計算兩者的軌跡誤差以顯示其吻合度,反映多視角生成的一致性。具體來說,我們計算絕對軌跡誤差 (Absolute Trajectory Error, ATE) 以及平移分量 (RPE-t) 與旋轉分量 (RPE-R) 的相對姿態誤差。為了公平比較,我們把軌跡長度正規化為 1.0,並排除相機移動極小的案例(如汽車停等紅燈)。
如表 25 所示,結果呼應多視角幾何一致性的發現:從 Cosmos WFM 微調而來的世界模型,其軌跡一致性遠優於 VideoLDM-MultiView。我們注意到,後訓練的 Cosmos 世界模型軌跡一致性已接近真實世界影片。
軌跡一致性:軌跡遵循誤差 (Trajectory Following Error, TFE)。 進一步地,對於接受軌跡控制條件輸入的模型,我們用上述相同的相機姿態估計流程,以多視角資訊計算前置相機的姿態,並把預測軌跡與真值軌跡條件比較,衡量模型遵循給定軌跡路徑的程度。如表 25 所示,用 Cosmos 後訓練世界模型生成影片估計的軌跡誤差,僅比真值基準低不到 7 公分的精度。如此微小的差距顯示模型能精確遵循給定的軌跡路徑,這對訓練自動駕駛代理至關重要。
物件追蹤一致性。 最後,我們用 YOLOv11x (Khanam and Hussain 2024) 對生成的 8 秒影片執行物件偵測與追蹤。人工標註者的任務是找出追蹤演算法誤判出物理上不可能情境的案例,例如兩個不同物體(如行人與汽車)錯誤地合併為單一追蹤實體。為此,我們提供標註者隨機抽樣的 20 部生成影片、共 157 個物體。值得注意的是,157 個物體中沒有任何一個出現物理上不可能的情境,展現了我們生成駕駛影片的物理一致性與物體恆存性。
7 護欄
為了安全使用我們的 WFM,我們開發了完善的護欄系統,包含兩個階段:pre-Guard 階段與 post-Guard 階段。pre-Guard 階段利用 Aegis (Ghosh et al. 2024) 與關鍵字清單阻擋有害提示詞;post-Guard 階段以影片內容安全分類器與人臉模糊過濾器阻擋有害的視覺輸出。流程如圖 30 所示。

圖 30:Cosmos 護欄概覽。 Cosmos 護欄包含 pre-Guard 與 post-Guard:pre-Guard 基於 Aegis (Ghosh et al. 2024) 與關鍵字阻擋輸入;post-Guard 基於影片內容安全分類器阻擋輸出,並對輸出的人臉進行模糊處理。
7.1 Pre-Guard
我們的 pre-Guard 是文字領域的護欄,由處理語意複雜提示詞的 LLM 護欄,以及針對明確不安全關鍵字的簡單黑名單檢查器組成。
7.1.1 關鍵字阻擋
黑名單啟發式規則是降低不安全內容生成風險的第一道防線。它針對明確有害的生成而設計,把提示詞與一份寫死的、包含大量露骨與令人反感詞彙的黑名單做關鍵字比對。輸入詞彙以 WordNetLemmatizer 做詞形還原 (Lemmatization)——這個工具利用英語詞彙資料庫 (Miller 1995) 從變化形中萃取詞根,例如「abacii」的詞根是「abacus」。還原後的詞彙再與黑名單中的詞彙比對,若發現任何不當字眼,整個提示詞即被拒絕。我們使用全面的關鍵字集合,最大程度保護使用者。
7.1.2 Aegis 護欄
作為第二道防線,我們使用 Aegis-AI-Content-Safety-LlamaGuard-LLM-Defensive-1.0 (Ghosh et al. 2024)——Llama-Guard (Inan et al. 2023) 的微調版本,在 NVIDIA 的 Aegis 內容安全資料集上訓練,涵蓋 NVIDIA 的 13 個關鍵安全風險類別的廣泛分類法。AEGIS 1.0 有防禦版與寬鬆版兩種版本,防禦版採用比寬鬆版更嚴格的許可邊界。Cosmos 使用防禦版 Aegis,阻擋企圖生成有害內容的潛在有害使用者提示詞。若輸入提示詞被此提示詞過濾器歸類為不安全,影片就不會生成,並顯示錯誤訊息。
使用 Aegis 作為提示詞過濾器時,落入下列類別的提示詞會被歸類為不安全:暴力、色情、犯罪策劃、武器、藥物濫用、自殺、兒少性虐待素材、仇恨、騷擾、威脅與粗俗言語。不落入上述類別的提示詞,從提示詞過濾的角度視為安全。
7.2 Post-Guard
我們的 post-Guard 是視覺領域的護欄,由影片內容安全過濾器與生成輸出的人臉模糊過濾器組成。
7.2.1 影片內容安全過濾器
影片內容安全過濾器是在我們的影片資料集與生成結果上訓練的影格層級多類別分類器。類別中有些視為安全、有些不安全。訓練此分類器的一大挑戰在於平衡偽陽性(安全內容被誤標為不安全)與偽陰性(不安全內容被誤判為安全)。為了把分類錯誤降到最低,我們在訓練時謹慎地平衡資料。
我們蒐集三種真值標註資料。第一,從資料集取樣大量影片、抽取影格,用 VLM 判定其類別。第二,用 WFM 以一組提示詞生成合成影片,確保涵蓋極端案例 (Corner Case) 與代表性最不足的內容類別。第三,人工標註者為部分資料集提供「黃金標準」標籤,加上關鍵的驗證層,幫助我們持續精進分類器的準確率。我們為每個影片影格萃取 SigLIP (Zhai et al. 2023) 嵌入,並在嵌入上訓練簡單的 MLP 分類器。推論時,我們為每個影格生成 SigLIP 嵌入再套用分類器;只要任一影格被判為不安全,整部影片即被標記為不安全。
7.2.2 人臉模糊過濾器
我們使用最先進的人臉偵測模型 RetinaFace (Deng et al. 2020),以高信心分數識別臉部區域。對任何大於 $20\times 20$ 像素的偵測到的人臉區域,我們施加像素化 (Pixelation) 遮蔽,同時保留整體場景構圖以供物理 AI 應用。
7.3 紅隊演練
我們設有專責的紅隊 (Red Team),使用蒐集於內部攻擊提示詞資料集的標準與對抗性 (Adversarial) 範例主動探測系統。這些影片輸出由一組為此任務特別訓練的專家標註者標註,依 7.1.2 節分類法的多個危害類別,以 1-5 分為生成影片評分。這些標註也會指明偵測到不安全內容的起訖影格,從而產生高品質標註。紅隊也用針對性範例獨立探測各個護欄元件,找出弱點並改進邊緣案例的表現。截至本文發表,紅隊已測試並標註超過 $10{,}000$ 組精心設計、涵蓋廣泛不安全內容的提示詞-影片配對。
8 相關工作
世界模型。 「世界模型」的概念源自 Ha and Schmidhuber 2018 的開創性工作,其提出以神經網路模型學習真實世界的表徵,在給定當前狀態與輸入下預測未來狀態。精確的物理世界模型表徵不僅能可靠地預測未來狀態,也支撐明智的決策。對物理世界建模的概念並不新穎;傳統自動化與機器人產業長期在規劃與控制演算法中使用基於物理定律與系統辨識 (System Identification) 的數學模型 (Murray et al. 2017)。然而這些系統專屬的模型通常侷限於低維狀態空間,限制了跨系統的泛化與知識轉移,也限制了模型在新任務或新環境中的重用。深度學習的近期進展、尤其是生成式 AI,讓直接從視覺觀察學習世界模型成為可能。
現代世界模型流程可依主幹架構分類。多數工作 (Hafner et al. 2019; Hafner et al. 2021; Kim et al. 2020; Kim et al. 2021; Hafner et al. 2023; Hansen et al. 2024),包括 Ha 與 Schmidhuber 的原始論文 (Ha and Schmidhuber 2018),採用循環神經網路 (Recurrent Neural Network) 在自編碼器學到的潛在空間中對系統狀態演化建模。較新的趨勢把世界模型視為視覺觀察空間中的生成模型,常以條件式影片生成模型的形式出現(如動作生成影片、文字生成影片)。這些模型可以是自迴歸式 (Yang et al. 2023; Micheli et al. 2023; Robine et al. 2023; Bruce et al. 2024; Liu et al. 2024b) 或基於擴散 (Valevski et al. 2024; Alonso et al. 2024; Ding et al. 2024),如同本文所考慮的。另一個有前景的方向是生成式模擬 (Generative Simulation) (Nasiriany et al. 2024; Hua et al. 2024),結合生成式 AI 與物理模擬器來建模真實世界。
訓練良好的世界模型有多種應用方式,包括驗證 (Hu et al. 2023)、基於規劃的模型預測控制 (Hansen et al. 2024; Bar et al. 2024),以及基於模型的強化學習 (Yang et al. 2023; Robine et al. 2023; Alonso et al. 2024; Zhang et al. 2024; Ding et al. 2024)。世界模型的有效性已在多個領域獲得證明,如電腦遊戲 (Hafner et al. 2021; Kim et al. 2020; Bruce et al. 2024; Valevski et al. 2024; Alonso et al. 2024)、真實世界機器人 (Wu et al. 2023b; Yang et al. 2023) 與自動駕駛 (Kim et al. 2021; Blattmann et al. 2023b; Hu et al. 2023; Zhao et al. 2024a)。我們預見基礎性的世界模型將對這些產業帶來變革性的影響。
影片生成模型。 影片生成模型領域近年發展迅速。從最初只能產生短小、低解析度影片的模型開始,此領域已顯著演進,影片生成模型如今站在生成式 AI 研究的最前沿 (Ho et al. 2022; Huang et al. 2024)。近年出現了 Sora、Dream Machine、Gen 3 與 Kling 等令人印象深刻的影片生成模型,能產生逼真的高解析度影片 (Luma 2024; OpenAI 2024b; KuaiShou 2024; Runway 2024)。這些進展距離第一個影片生成模型的發布不過短短數年。
影片生成模型的既有工作多聚焦於文字生成影片任務,根據文字提示詞輸入生成影片 (Yang et al. 2024d; Ma et al. 2024; Lin et al. 2024a; Blattmann et al. 2023a; Ge et al. 2023; Girdhar et al. 2024)。這些模型讓使用者能用精心設計的文字提示詞創作出色的影片。其他熱門任務包括從給定圖像影格開始生成影片的圖像生成影片 (Image-to-Video) (Blattmann et al. 2023a; Wang et al. 2024c; Ren et al. 2024; Wang et al. 2021b; Mallya et al. 2022; Gururani et al. 2023)、根據參考影片生成新影片的影片生成影片 (Video-to-Video) (Wang et al. 2018; Wang et al. 2019; Mallya et al. 2020; Ku et al. 2024; Liu et al. 2024a),以及隨著世界模型與具身 AI (Embodied AI) 發展而興起、根據動作生成影片的動作生成影片 (Action-to-Video) (Bruce et al. 2024; Valevski et al. 2024; Alonso et al. 2024; Tulyakov et al. 2018)。
大多數影片生成模型採用擴散模型框架 (Blattmann et al. 2023a; Lin et al. 2024a; Ge et al. 2023; Ma et al. 2024; Yang et al. 2024d),把雜訊逐步轉化為影片序列。自迴歸模型也被用於影片生成,優點是能以統一方式處理影片與其他模態 (Kondratyuk et al. 2024; Deng et al. 2024; Liu et al. 2024c)。雖然自迴歸模型展現潛力,基於擴散的影片模型在視覺品質上仍更勝一籌。我們的目標是幫助物理 AI 開發者推進其應用。我們認為基於擴散與基於自迴歸的模型各有優劣:擴散模型能渲染視覺品質更好的影片;自迴歸模型能更好地利用 LLM 社群開發的各種技術。我們同時建立基於擴散(Cosmos-Diffusion)與基於自迴歸(Cosmos-Autoregressive)的 WFM,並提供給物理 AI 開發者。
具相機控制的影片生成。 3D 一致的影片生成可追溯至視角合成與 3D 重建的早期工作,當時社群試圖以套用於各種 3D 表徵的神經渲染 (Neural Rendering) 建立 3D 一致的影片 (Zhou et al. 2018; Mildenhall et al. 2020; Wang et al. 2021a; Li et al. 2023; Kerbl et al. 2023)。在這條研究線中,單張圖像的 3D 視角合成 (Tucker and Snavely 2020; Wiles et al. 2020; Yu et al. 2021; Lin et al. 2023b; Charatan et al. 2024) 特別具挑戰性,通常需要從多視角圖像資料集學習強力的 3D 先驗模型。由於這類 3D 先驗模型往往難以規模化,學習式視角合成也被以純資料驅動的方式、使用可規模化的 Transformer 架構 (Vaswani et al. 2017; Dosovitskiy et al. 2021) 進行探索。這繞過了對顯式 3D 先驗知識的需求 (Rombach et al. 2021; Sajjadi et al. 2022):不依賴套用於 3D 表徵的神經渲染,而是由以相機輸入為條件的神經網路直接合成新視角 (Tatarchenko et al. 2016)。此範式已成功以擴散模型規模化 (Liu et al. 2023b),在 3D 資產生成中找到廣泛應用 (Poole et al. 2023; Lin et al. 2023a; Shi et al. 2023; Qian et al. 2024; Li et al. 2024; NVIDIA 2024d)。影片生成品質的近期進展顯示,透過擴大訓練影片資料,有潛力達成完整的 3D 一致性 (Brooks et al. 2024)。此類模型的相機可控性因其在機器人與自主導航上的巨大應用潛力,已成為活躍的研究領域 (He et al. 2024a; Wang et al. 2024g; Xu et al. 2024)。
機器人控制的生成模型。 深度生成模型的近期進展引發了將其應用於機器人控制的濃厚興趣。已出現數種方法:一條研究線直接把擴散模型用作視覺運動策略 (Visuomotor Policy),在多種機器人任務的模仿學習 (Imitation Learning) 上展現大幅改進 (Chi et al. 2023; Wang et al. 2024f; Prasad et al. 2024; Ke et al. 2024)。與本工作更相關的另外兩條線是:把預訓練圖像與影片生成模型用作運動規劃器,以及利用圖像與影片資料做生成式預訓練。生成式運動規劃方法 (Ko et al. 2024; Zhou et al. 2024b; Finn and Levine 2017; Black et al. 2023; Du et al. 2024) 透過生成中間視覺子目標 (Sub-goal) 而非顯式動作序列,來增強對未見環境的泛化。這種視覺表徵策略更為穩健,因為圖像與影片子目標能跨越多樣的環境配置泛化,不像動作序列通常侷限於特定環境與任務。生成式預訓練方法 (Gupta et al. 2024b; Cheang et al. 2024; He et al. 2024b) 利用大規模圖像與影片資料集做預訓練。Gupta et al. 2024b 從預訓練文字生成圖像擴散模型中萃取並利用特徵來引導後續的策略學習;Cheang et al. 2024 與 He et al. 2024b 則採用兩階段框架:先預訓練模型預測未來影格,再微調模型聯合預測動作與未來影格。
自動駕駛的生成模型。 影片生成模型有潛力革新自動駕駛模擬,讓系統能以多樣的輸入模態(如文字、圖像、軌跡、3D 資料或地圖)為條件生成逼真的駕駛影片 (Kim et al. 2021; Blattmann et al. 2023b; Gao et al. 2024d; Wang et al. 2023a; Wang et al. 2024e; Lu et al. 2025; Jia et al. 2023; Yang et al. 2024c; Hu et al. 2023; Gao et al. 2024c; Gao et al. 2024b)。儘管潛力可觀,既有方法受限於資料規模 (Wang et al. 2023a; Wang et al. 2024e; Lu et al. 2025; Jia et al. 2023; Gao et al. 2024c; Gao et al. 2024b)、解析度 (Yang et al. 2024c; Hu et al. 2023) 與相機視角數 (Blattmann et al. 2023b; Gao et al. 2024d),限制了其作為全面駕駛世界模擬器的效果。為了克服這些限制,我們利用強大的預訓練 WFM 的能力,開發彈性且可規模化的駕駛模擬器。我們的模型達到高解析度、高影格率與多視角一致性。
Tokenizer。 學習能重現輸入視覺資料的潛在特徵已有相當長的歷史 (Kingma 2013; van den Oord et al. 2017; Hinton et al. 1995; He et al. 2022)。近年來,這類模型(也稱 tokenizer)已被廣泛納入為關鍵元件,用以提升大規模生成模型的訓練效率 (Rombach et al. 2022; Esser et al. 2021)。
連續視覺 tokenizer——通常包括自編碼器 (AE) 與變分自編碼器 (VAE)——把視覺資料壓縮到連續潛在空間,讓基於擴散的模型能在其上高效訓練 (Song et al. 2020; Lipman et al. 2022; Ho et al. 2020)。推論時,生成的潛在表徵再由 tokenizer 解碼器解碼回 RGB 空間。各種擴散模型都以這種方式訓練,用於圖像 (Rombach et al. 2022; Ramesh et al. 2022; Betker et al. 2023; Dai et al. 2023; Podell et al. 2024; FLUX 2024; Gafni et al. 2022) 與影片生成 (Zeng et al. 2024; Blattmann et al. 2023b; Blattmann et al. 2023a; Ge et al. 2023; Brooks et al. 2024; Wang et al. 2023b; Yu et al. 2023b; An et al. 2023; Girdhar et al. 2024)。
離散視覺 tokenizer 額外包含一個量化器 (Quantizer) (van den Oord et al. 2017; Zhao et al. 2024b; Mentzer et al. 2023; Yu et al. 2024a; Lee et al. 2022; Yu et al. 2024b),把連續潛在表徵進一步離散化到離散空間,便於與文字、音訊等其他模態一起整合進大型語言模型與視覺語言模型。因此,離散 tokenizer 被部署於各種視覺理解 (Wu et al. 2024; Team 2024a; Sun et al. 2024b; Wang et al. 2024d) 以及圖像 (Esser et al. 2021; Ramesh et al. 2021; Yu et al. 2022; Chang et al. 2022; Sun et al. 2024a) 與影片生成任務 (Yan et al. 2021; Villegas et al. 2023; Hong et al. 2023; Wu et al. 2022; Ge et al. 2022; Yu et al. 2023a; Luo et al. 2024; Kondratyuk et al. 2024)。
Cosmos tokenizer 廣泛建立在先前研究的基礎上,例如 FSQ (Mentzer et al. 2023) 與因果架構 (Yu et al. 2023a),目標是打造一套高效、高品質的 tokenizer。
9 結論與討論
Cosmos 世界基礎模型是朝向建立物理世界通用模擬器的重要一步。本文概述了我們的完整方法,包括資料整理流程、連續與離散 tokenizer 的設計、擴散與自迴歸世界基礎模型的架構,以及面向多樣下游物理 AI 任務的微調過程。值得注意的是,我們展示了預訓練世界模型對關鍵應用的適應性,包括 3D 世界導覽、機器人操作與自動駕駛系統——這些應用同時要求 3D 一致性與動作可控性。
侷限。 儘管有所進展,世界基礎模型的發展仍處於早期階段。目前的模型(包括我們的)尚不足以作為物理世界的可靠模擬器。我們觀察到模型仍存在問題,包括缺乏物體恆存性、接觸密集動態 (Contact-rich Dynamics) 的不精確,以及指令遵循的不一致。此外,生成影片的逼真度並不總是等同於對基本物理原理的遵循,例如重力、光的交互作用與流體動力學。
評估是另一項重大挑戰。為人類評估物理擬真度定義穩健的評分準則很困難,因為這類評估常受個人偏見、背景與其他主觀因素影響。而且這些評估與下游物理 AI 任務使用的指標未必正相關。為了應對這些挑戰,有前景的方向包括開發由多模態 LLM 驅動的自動評估器,以及利用既有物理模擬器實現可重現、可互動的評估,從而減少對人工評估的依賴。
自迴歸 vs 擴散 WFM。 我們在 3D 一致性(5.3.1 節)與機器人影片生成(6.2 節)的評估結果顯示,基於擴散的 WFM 目前提供更好的生成品質。透過微調,基於擴散的 WFM 能納入多樣的控制訊號,包括相機姿態、末端執行器 (End-effector) 位置或自動駕駛車輛軌跡,並生成多視角影片等新格式的輸出。然而,基於自迴歸的 WFM 蘊含尚未開發的巨大潛力:它們可以 (1) 利用大型語言模型的預訓練權重繼承廣泛的世界知識,(2) 藉助為因果注意力設計的先進推論最佳化技術實現更快的生成。若這些能力充分實現,自迴歸 WFM 可能特別適合需要互動控制或即時處理的應用,例如機器人領域的規劃與模擬。重要的是,擴散與自迴歸模型之間的界線並非壁壘分明。近期進展顯示,具雙向注意力的擴散 transformer 可以蒸餾為具因果注意力的學生 transformer,從而在推論時支援 key-value 快取 (Yin et al. 2024);同樣地,自迴歸模型可以納入局部雙向注意力,經由擴散頭生成圖像 (Zhou et al. 2024a)。探索這些混合方法及其權衡仍是活躍且有前景的研究領域。我們計畫進一步研究這些形式,並在未來工作中提供全面的分析。
附錄 A 貢獻者與致謝
A.1 核心貢獻者
- 資料整理: Jacob Huffman, Francesco Ferroni, Alice Luo, Niket Agarwal, Hao Wang, Jing Zhang, David Page, Vasanth Rao Naik Sabavat, Sriharsha Niverty, Erik Barker, Lindsey Pavao, Stella Shi, Prithvijit Chattopadhyay, Shitao Tang, Yin Cui, Yunhao Ge, Qianli Ma, Yifan Ding, Seungjun Nah, Siddharth Gururani, Jiashu Xu, Grace Lam, Tiffany Cai, Jibin Varghese, Pooya Jannaty, Jay Zhangjie Wu, Yuxuan Zhang, Huan Ling, Hanzi Mao, Heng Wang
- Tokenizer: Jinwei Gu, Xian Liu, Songwei Ge, Ting-Chun Wang, Haoxiang Wang, Fitsum Reda
- 基於擴散的世界基礎模型預訓練: Qinsheng Zhang, Lin Yen-Chen, Xiaohui Zeng, Huan Ling, Shitao Tang, Maciej Bala, Ting-Chun Wang, Yu Zeng, Seungjun Nah, Qianli Ma, Hanzi Mao
- 基於自迴歸的世界基礎模型預訓練: Haoxiang Wang, Yifan Ding, Xian Liu, Jiaojiao Fan, Xiaohui Zeng, Yogesh Balaji
- 提示詞上採樣器: Yunhao Ge, Haoxiang Wang, Jiashu Xu, Yin Cui
- 擴散解碼器: Huan Ling, Jiaojiao Fan, Fitsum Reda, Yogesh Balaji, Hanzi Mao, Qinsheng Zhang
- 3D 一致性預訓練評估: Jiahui Huang, Chen-Hsuan Lin
- 物理對齊預訓練評估: Francesco Ferroni, Prithvijit Chattopadhyay, Xinyue Wei, Qianli Ma, Gergely Klár, Chen-Hsuan Lin
- 相機控制後訓練評估: Xiaohui Zeng, Tsung-Yi Lin, Jingyi Jin, Chen-Hsuan Lin
- 機器人後訓練評估: Lin Yen-Chen, Wei-Cheng Tseng, Yunhao Ge, Xian Liu, Shitao Tang, Fangyin Wei, Lyne Tchapmi, Yu Zeng, Qingqing Zhao, Yin Cui, Zhaoshuo Li, Jinwei Gu
- 自動駕駛後訓練評估: Seung Wook Kim, Jay Zhangjie Wu, Jiahui Huang, Francesco Ferroni, Michele Fenzi, Daniel Dworakowski, Despoina Paschalidou, Ed Schmerling, Shiyi Lan, Laura Leal-Taixe, Sanja Fidler, Huan Ling
- 護欄: Jibin Varghese, Arslan Ali, Grace Lam, Pooya Jannaty
- 平台架構師: Ming-Yu Liu
A.2 貢獻者
Anqi Li, Arsalan Mousavian, Artur Zolkowski, Bartosz Stefaniak, Dieter Fox, Ethan He, Kaichun Mo, Morteza Ramezanali, Przemek Tredak, Wei Yang, Xiaowei Ren, Yongxin Chen, Zeeshan Patel
A.3 致謝
我們感謝 1X Technologies 慷慨提供人形機器人資料,並為本技術報告中機器人操作的後訓練提供寶貴支持。
我們感謝 Aarti Basant, Akan Huang, Alex Qi, Alexis Bjorlin, Amanda Moran, Amol Fasale, Ankit Patel, Arash Vahdat, Aryaman Gupta, Ashna Khetan, Ashwath Aithal, Bor-Yiing Su, Bryan Catanzaro, Charles Hsu, Chris Pruett, Christopher Horvath, Clark Doan, Coulten Holt, Dane Aconfora, Deepak Narayanan, Dennis Chang, Dheeraj Kapur, Dong Ahn, Ebrar Erdem, Elmar Haussmann, Fuzhao Xue, Gandhi Vaithilingam, Henry Estela, Henry Vera, Herb Woodruff, Imad El Hanafi, Jashojit Mukherjee, Jason Sewall, Jensen Huang, John Dickinson, Jonah Alben, Jonah Philion, Josh Abbott, Jun Gao, Kumar Anik, Lee Ditiangkin, Ligeng Zhu, Linxi Fan, Luke Alonso, Madison Huang, Marek Dabek, Mark Arnold, Max Ehrlich, Michele Ferretti, Misbah Mubarak, Misha Smelyanskiy, Mohamed Fawzy, Mohammad Harrim, Mohammad Shoeybi, Omkar Mehta, Pallab Bhattacharya, Paniz Karbasi, Pasha Shamis, Raju Wagwani, Rick Izzo, Robert Hero, Sharon Clay, Song Han, Songyan Tang, Sophia Huang, Sridhar Bhuvanapalli, TJ Galda, Thomas Volk, Tobias Lasser, Vaibhav Ranglani, Vijay Anand Korthikanti, Yao Lu, Yazdan Aghaghiri, Yugi Guvvala, Yuke Zhu 與 Zekun Hao 的回饋與工程支援。
我們感謝 Iain Cunningham, Jim Fan, Marco Pavone, Meredith Price, Nikki Pope 與 Scott Reed 對本技術報告初稿的回饋。
參考文獻
- Abbas et al. (2023) Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Morcos. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540, 2023.
- Adler et al. (2024) Bo Adler, Niket Agarwal, Ashwath Aithal, Dong H Anh, Pallab Bhattacharya, Annika Brundyn, Jared Casper, Bryan Catanzaro, Sharon Clay, Jonathan Cohen, et al. Nemotron-4 340b technical report. arXiv preprint arXiv:2406.11704, 2024.
- Agrawal et al. (2024) Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Baptiste Bout, Devendra Chaplot, Jessica Chudnovsky, Diogo Costa, Baudouin De Monicault, Saurabh Garg, Theophile Gervet, et al. Pixtral 12b. arXiv preprint arXiv:2410.07073, 2024.
- AI Image Lab, University of Modena (2016) AI Image Lab, University of Modena. Bbc planet earth dataset, 2016. URL https://aimagelab.ing.unimore.it/imagelab/researchActivity.asp?idActivity=19. Accessed: 2024-10-17.
- Alonso et al. (2024) Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos Storkey, Tim Pearce, and François Fleuret. Diffusion for world modeling: Visual details matter in atari. In NeurIPS, 2024.
- An et al. (2023) Jie An, Songyang Zhang, Harry Yang, Sonal Gupta, Jia-Bin Huang, Jiebo Luo, and Xi Yin. Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv preprint arXiv:2304.08477, 2023.
- Atzmon et al. (2024) Yuval Atzmon, Maciej Bala, Yogesh Balaji, Tiffany Cai, Yin Cui, Jiaojiao Fan, Yunhao Ge, Siddharth Gururani, Jacob Huffman, Ronald Isaac, et al. Edify image: High-quality image generation with pixel space laplacian diffusion models. arXiv preprint arXiv:2411.07126, 2024.
- Balaji et al. (2022) Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022.
- Bar et al. (2024) Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models. arXiv preprint arXiv:2412.03572, 2024.
- Betker et al. (2023) James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2023.
- Black et al. (2023) Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Walke, Chelsea Finn, Aviral Kumar, and Sergey Levine. Zero-shot robotic manipulation with pre-trained image-editing diffusion models. In NeurIPS Workshops, 2023.
- Blattmann et al. (2023a) Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023a.
- Blattmann et al. (2023b) Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In CVPR, 2023b.
- Brooks et al. (2024) Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators, 2024. URL https://openai.com/research/video-generation-models-as-world-simulators.
- Brown et al. (2020) Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020.
- Bruce et al. (2024) Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. In ICML, 2024.
- Cai et al. (2024) Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. In ICML, 2024.
- Castellano (2024) Brandon Castellano. Pyscenedetect, 2024. URL https://www.scenedetect.com. Video Cut Detection and Analysis Tool.
- Chang et al. (2022) Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. Maskgit: Masked generative image transformer. In CVPR, 2022.
- Charatan et al. (2024) David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In CVPR, 2024.
- Cheang et al. (2024) Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al. Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation. arXiv preprint arXiv:2410.06158, 2024.
- Chen et al. (2016) Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174, 2016.
- Chen (2023) Ting Chen. On the importance of noise scheduling for diffusion models. arXiv preprint arXiv:2301.10972, 2023.
- Chen et al. (2024) Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In CVPR, 2024.
- Chi et al. (2023) Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. RSS, 2023.
- Dai et al. (2023) Xiaoliang Dai, Ji Hou, Chih-Yao Ma, Sam Tsai, Jialiang Wang, Rui Wang, Peizhao Zhang, Simon Vandenhende, Xiaofang Wang, Abhimanyu Dubey, et al. Emu: Enhancing image generation models using photogenic needles in a haystack. arXiv preprint arXiv:2309.15807, 2023.
- de Brébisson and Vincent (2016) Alexandre de Brébisson and Pascal Vincent. The z-loss: a shift and scale invariant classification loss belonging to the spherical family. arXiv preprint arXiv:1604.08859, 2016.
- Dehghani et al. (2023) Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Peter Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. In ICML, 2023.
- Deng et al. (2024) Haoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi, and Xinlong Wang. Autoregressive video generation without vector quantization. arXiv preprint arXiv:2412.14169, 2024.
- Deng et al. (2009) Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
- Deng et al. (2020) Jiankang Deng, Jia Guo, Yuxiang Zhou, Jinke Yu, Irene Kotsia, and Stefanos Zafeiriou. Retinaface: Single-stage dense face localisation in the wild. CVPR, 2020.
- DeTone et al. (2018) Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. Superpoint: Self-supervised interest point detection and description. In CVPR Workshops, 2018.
- Ding et al. (2024) Zihan Ding, Amy Zhang, Yuandong Tian, and Qinqing Zheng. Diffusion world model: Future modeling beyond step-by-step rollout for offline reinforcement learning. arXiv preprint arXiv:2402.03570, 2024.
- Dosovitskiy et al. (2021) Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: transformers for image recognition at scale. In ICLR, 2021.
- Du et al. (2024) Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel. Learning universal policies via text-guided video generation. In NeurIPS, 2024.
- Dubey et al. (2024) Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024.
- Ebert et al. (2022) Frederik Ebert, Yanlai Yang, Karl Schmeckpeper, Bernadette Bucher, Georgios Georgakis, Kostas Daniilidis, Chelsea Finn, and Sergey Levine. Bridge data: Boosting generalization of robotic skills with cross-domain datasets. In RSS, 2022.
- Esser et al. (2021) Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021.
- Esser et al. (2024) Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling rectified flow transformers for high-resolution image synthesis. In ICML, 2024.
- Farnebäck (2003) Gunnar Farnebäck. Two-frame motion estimation based on polynomial expansion. In Scandinavian Conference on Image Analysis, 2003.
- Finn and Levine (2017) Chelsea Finn and Sergey Levine. Deep visual foresight for planning robot motion. In ICRA, 2017.
- FLUX (2024) FLUX. FLUX.1: Image generation, 2024. URL https://huggingface.co/black-forest-labs/FLUX.1-dev.
- Fu et al. (2023) Stephanie Fu, Netanel Yakir Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dreamsim: Learning new dimensions of human visual similarity using synthetic data. In NeurIPS, 2023.
- Gadre et al. (2024) Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, et al. Datacomp: In search of the next generation of multimodal datasets. In NeurIPS, 2024.
- Gafni et al. (2022) Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. In ECCV, 2022.
- Gage (1994) Philip Gage. A new algorithm for data compression. The C Users Journal, 1994.
- Gao et al. (2024a) Ruiqi Gao, Emiel Hoogeboom, Jonathan Heek, Valentin De Bortoli, Kevin P. Murphy, and Tim Salimans. Diffusion meets flow matching: Two sides of the same coin, 2024a. URL https://diffusionflow.github.io/.
- Gao et al. (2024b) Ruiyuan Gao, Kai Chen, Bo Xiao, Lanqing Hong, Zhenguo Li, and Qiang Xu. Magicdrivedit: High-resolution long video generation for autonomous driving with adaptive control. arXiv preprint arXiv:2411.13807, 2024b.
- Gao et al. (2024c) Ruiyuan Gao, Kai Chen, Enze Xie, Lanqing HONG, Zhenguo Li, Dit-Yan Yeung, and Qiang Xu. Magicdrive: Street view generation with diverse 3d geometry control. In ICLR, 2024c.
- Gao et al. (2024d) Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, and Hongyang Li. Vista: A generalizable driving world model with high fidelity and versatile controllability. In NeurIPS, 2024d.
- Gatys et al. (2016) Leon A Gatys, Alexander S Ecker, and Matthias Bethge. Image style transfer using convolutional neural networks. In CVPR, 2016.
- Ge et al. (2022) Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin Huang, and Devi Parikh. Long video generation with time-agnostic vqgan and time-sensitive transformer. In ECCV, 2022.
- Ge et al. (2023) Songwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon, Andrew Tao, Bryan Catanzaro, David Jacobs, Jia-Bin Huang, Ming-Yu Liu, and Yogesh Balaji. Preserve your own correlation: A noise prior for video diffusion models. In ICCV, 2023.
- Ge et al. (2024) Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman, Tsung-Yi Lin, Ming-Yu Liu, and Yin Cui. Visual fact checker: Enabling high-fidelity detailed caption generation. In CVPR, 2024.
- Ghosh et al. (2024) Shaona Ghosh, Prasoon Varshney, Erick Galinkin, and Christopher Parisien. Aegis: Online adaptive ai content safety moderation with ensemble of llm experts. arXiv preprint arXiv:2404.05993, 2024.
- Girdhar et al. (2023) Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. Imagebind: One embedding space to bind them all. In CVPR, 2023.
- Girdhar et al. (2024) Rohit Girdhar, Mannat Singh, Andrew Brown, Quentin Duval, Samaneh Azadi, Sai Saketh Rambhatla, Akbar Shah, Xi Yin, Devi Parikh, and Ishan Misra. Emu video: Factorizing text-to-video generation by explicit image conditioning. In ECCV, 2024.
- Grauman et al. (2024) Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In CVPR, 2024.
- Gupta et al. (2024a) Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and José Lezama. Photorealistic video generation with diffusion models. In ECCV, 2024a.
- Gupta et al. (2024b) Gunshi Gupta, Karmesh Yadav, Yarin Gal, Dhruv Batra, Zsolt Kira, Cong Lu, and Tim GJ Rudner. Pre-trained text-to-image diffusion models are versatile representation learners for control. In ICLR Workshops, 2024b.
- Gururani et al. (2023) Siddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle, and Ming-Yu Liu. SPACE: Speech-driven Portrait Animation with Controllable Expression. In ICCV, 2023.
- Ha and Schmidhuber (2018) David Ha and Jürgen Schmidhuber. World models. arXiv preprint arXiv:1803.10122, 2018.
- Hafner et al. (2019) Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:1912.01603, 2019.
- Hafner et al. (2021) Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. In ICLR, 2021.
- Hafner et al. (2023) Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023.
- Hansen et al. (2024) Nicklas Hansen, Hao Su, and Xiaolong Wang. Td-mpc2: Scalable, robust world models for continuous control. In ICLR, 2024.
- Hartley and Zisserman (2003) Richard Hartley and Andrew Zisserman. Multiple view geometry in computer vision. Cambridge university press, 2003.
- He et al. (2024a) Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. Cameractrl: Enabling camera control for text-to-video generation. arXiv preprint arXiv:2404.02101, 2024a.
- He et al. (2024b) Haoran He, Chenjia Bai, Ling Pan, Weinan Zhang, Bin Zhao, and Xuelong Li. Learning an actionable discrete diffusion policy via large-scale actionless video pre-training. In NeurIPS, 2024b.
- He et al. (2022) Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In CVPR, 2022.
- Heusel et al. (2017) Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurIPS, 2017.
- Hinton et al. (1995) Geoffrey E Hinton, Peter Dayan, Brendan J Frey, and Radford M Neal. The "wake-sleep" algorithm for unsupervised neural networks. Science, 1995.
- Ho and Salimans (2022) Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
- Ho et al. (2020) Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020.
- Ho et al. (2022) Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In NeurIPS, 2022.
- Hong et al. (2023) Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In ICLR, 2023.
- Hoogeboom et al. (2023) Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. Simple diffusion: End-to-end diffusion for high resolution images. In ICML, 2023.
- Hoogeboom et al. (2024) Emiel Hoogeboom, Thomas Mensink, Jonathan Heek, Kay Lamerigts, Ruiqi Gao, and Tim Salimans. Simpler diffusion (sid2): 1.5 fid on imagenet512 with pixel-space diffusion. arXiv preprint arXiv:2410.19324, 2024.
- Hu et al. (2023) Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. Gaia-1: A generative world model for autonomous driving. arXiv preprint arXiv:2309.17080, 2023.
- Hu et al. (2022) Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In ICLR, 2022.
- Hua et al. (2024) Pu Hua, Minghuan Liu, Annabella Macaluso, Yunfeng Lin, Weinan Zhang, Huazhe Xu, and Lirui Wang. Gensim2: Scaling robot data generation with multi-modal and reasoning llms. arXiv preprint arXiv:2410.03645, 2024.
- Huang et al. (2024) Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. In CVPR, 2024.
- Inan et al. (2023) Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023.
- Jacobs et al. (2023) Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, and Yuxiong He. Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models. arXiv preprint arXiv:2309.14509, 2023.
- Jia et al. (2023) Fan Jia, Weixin Mao, Yingfei Liu, Yucheng Zhao, Yuqing Wen, Chi Zhang, Xiangyu Zhang, and Tiancai Wang. Adriver-i: A general world model for autonomous driving. arXiv preprint arXiv:2311.13549, 2023.
- Jiang et al. (2023) Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
- Kang et al. (2024) Bingyi Kang, Yang Yue, Rui Lu, Zhijie Lin, Yang Zhao, Kaixin Wang, Gao Huang, and Jiashi Feng. How far is video generation from world model? – a physical law perspective. arXiv preprint arXiv:2411.02385, 2024.
- Kaplan et al. (2020) Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
- Karras et al. (2020) Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In CVPR, 2020.
- Karras et al. (2022) Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In NeurIPS, 2022.
- Karras et al. (2024) Tero Karras, Miika Aittala, Jaakko Lehtinen, Janne Hellsten, Timo Aila, and Samuli Laine. Analyzing and improving the training dynamics of diffusion models. In CVPR, 2024.
- Ke et al. (2024) Tsung-Wei Ke, Nikolaos Gkanatsios, and Katerina Fragkiadaki. 3d diffuser actor: Policy diffusion with 3d scene representations. In CoRL, 2024.
- Kerbl et al. (2023) Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (TOG), 2023.
- Khanam and Hussain (2024) Rahima Khanam and Muhammad Hussain. Yolov11: An overview of the key architectural enhancements. arXiv preprint arXiv:2410.17725, 2024.
- Kim et al. (2024) Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024.
- Kim et al. (2020) Seung Wook Kim, Yuhao Zhou, Jonah Philion, Antonio Torralba, and Sanja Fidler. Learning to Simulate Dynamic Environments with GameGAN. In CVPR, 2020.
- Kim et al. (2021) Seung Wook Kim, , Jonah Philion, Antonio Torralba, and Sanja Fidler. DriveGAN: Towards a Controllable High-Quality Neural Simulation. In CVPR, 2021.
- Kingma (2013) Diederik P Kingma. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
- Ko et al. (2024) Po-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun, and Joshua B Tenenbaum. Learning to act from actionless videos through dense correspondences. In ICLR, 2024.
- Kondratyuk et al. (2024) Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, et al. Videopoet: A large language model for zero-shot video generation. In ICML, 2024.
- Kong et al. (2024) Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models. arXiv preprint arXiv:2412.03603, 2024.
- Korthikanti et al. (2023) Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems, 2023.
- Ku et al. (2024) Max Ku, Cong Wei, Weiming Ren, Huan Yang, and Wenhu Chen. Anyv2v: A tuning-free framework for any video-to-video editing tasks. TMLR, 2024.
- KuaiShou (2024) KuaiShou. Kling, 2024. URL https://klingai.com/.
- Lee et al. (2022) Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han. Autoregressive image generation using residual quantization. In CVPR, 2022.
- Lei Ba et al. (2016) Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
- Leviathan et al. (2023) Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In ICML, 2023.
- Li et al. (2024) Jiahao Li, Hao Tan, Kai Zhang, Zexiang Xu, Fujun Luan, Yinghao Xu, Yicong Hong, Kalyan Sunkavalli, Greg Shakhnarovich, and Sai Bi. Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model. In ICLR, 2024.
- Li et al. (2023) Zhaoshuo Li, Thomas Müller, Alex Evans, Russell H Taylor, Mathias Unberath, Ming-Yu Liu, and Chen-Hsuan Lin. Neuralangelo: High-fidelity neural surface reconstruction. In CVPR, 2023.
- Liang et al. (2024) Hanxue Liang, Jiawei Ren, Ashkan Mirzaei, Antonio Torralba, Ziwei Liu, Igor Gilitschenski, Sanja Fidler, Cengiz Oztireli, Huan Ling, Zan Gojcic, et al. Feed-forward bullet-time reconstruction of dynamic scenes from monocular videos. arXiv preprint arXiv:2412.03526, 2024.
- Lin et al. (2024a) Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Zhu, Shaodong Wang, Xianyi He, Yang Ye, Shenghai Yuan, Liuhan Chen, et al. Open-sora plan: Open-source large video generation model. arXiv preprint arXiv:2412.00131, 2024a.
- Lin et al. (2021) Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey. Barf: Bundle-adjusting neural radiance fields. In ICCV, 2021.
- Lin et al. (2023a) Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023a.
- Lin et al. (2024b) Ji Lin, Hongxu Yin, Wei Ping, Pavlo Molchanov, Mohammad Shoeybi, and Song Han. Vila: On pre-training for visual language models. In CVPR, 2024b.
- Lin et al. (2023b) Kai-En Lin, Yen-Chen Lin, Wei-Sheng Lai, Tsung-Yi Lin, Yi-Chang Shih, and Ravi Ramamoorthi. Vision transformer for nerf-based view synthesis from a single input image. In WACV, 2023b.
- Lin et al. (2014) Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
- Lindenberger et al. (2023) Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Pollefeys. Lightglue: Local feature matching at light speed. In ICCV, 2023.
- Ling et al. (2024) Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision. In CVPR, 2024.
- Lipman et al. (2022) Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.
- Liu et al. (2024a) Chang Liu, Rui Li, Kaidong Zhang, Yunwei Lan, and Dong Liu. Stablev2v: Stablizing shape consistency in video-to-video editing. arXiv preprint arXiv:2411.11045, 2024a.
- Liu et al. (2023a) Hao Liu, Matei Zaharia, and Pieter Abbeel. Ring attention with blockwise transformers for near-infinite context. arXiv preprint arXiv:2310.01889, 2023a.
- Liu et al. (2024b) Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. World model on million-length video and language with blockwise ringattention. CoRR, 2024b.
- Liu et al. (2024c) Haozhe Liu, Shikun Liu, Zijian Zhou, Mengmeng Xu, Yanping Xie, Xiao Han, Juan C Pérez, Ding Liu, Kumara Kahatapitiya, Menglin Jia, et al. Mardini: Masked autoregressive diffusion for video generation at scale. arXiv preprint arXiv:2410.20280, 2024c.
- Liu et al. (2023b) Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In ICCV, 2023b.
- Loshchilov and Hutter (2019) Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019.
- Lu et al. (2025) Jiachen Lu, Ze Huang, Zeyu Yang, Jiahui Zhang, and Li Zhang. Wovogen: World volume-aware diffusion for controllable multi-camera driving scene generation. In ECCV, 2025.
- Luma (2024) Luma. Dream machine, 2024. URL https://lumalabs.ai/dream-machine.
- Luo et al. (2024) Zhuoyan Luo, Fengyuan Shi, Yixiao Ge, Yujiu Yang, Limin Wang, and Ying Shan. Open-magvit2: An open-source project toward democratizing auto-regressive visual generation. arXiv preprint arXiv:2409.04410, 2024.
- Ma et al. (2024) Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. Latte: Latent diffusion transformer for video generation. arXiv preprint arXiv:2401.03048, 2024.
- Mallya et al. (2020) Arun Mallya, Ting-Chun Wang, Karan Sapra, and Ming-Yu Liu. World-consistent video-to-video synthesis. In ECCV, 2020.
- Mallya et al. (2022) Arun Mallya, Ting-Chun Wang, and Ming-Yu Liu. Implicit Warping for Animation with Image Sets. In NeurIPS, 2022.
- Mentzer et al. (2023) Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. Finite scalar quantization: Vq-vae made simple. arXiv preprint arXiv:2309.15505, 2023.
- Micheli et al. (2023) Vincent Micheli, Eloi Alonso, and François Fleuret. Transformers are sample-efficient world models. In ICLR, 2023.
- Mildenhall et al. (2020) Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
- Miller (1995) George A Miller. Wordnet: a lexical database for english. Communications of the ACM, 1995.
- Mistral and NVIDIA (2024) Mistral and NVIDIA. Mistral-nemo-12b-instruct: A 12b parameter large language model, 2024. URL https://mistral.ai/news/mistral-nemo/.
- Moritz et al. (2017) Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, William Paul, Michael I. Jordan, and Ion Stoica. Ray: A distributed framework for emerging AI applications. CoRR, abs/1712.05889, 2017. URL http://arxiv.org/abs/1712.05889.
- Murray et al. (2017) Richard M Murray, Zexiang Li, and S Shankar Sastry. A mathematical introduction to robotic manipulation. CRC press, 2017.
- Nasiriany et al. (2024) Soroush Nasiriany, Abhiram Maddukuri, Lance Zhang, Adeet Parikh, Aaron Lo, Abhishek Joshi, Ajay Mandlekar, and Yuke Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523, 2024.
- NVIDIA (2024a) NVIDIA. Isaac sim, 2024a. URL https://developer.nvidia.com/isaac/sim.
- NVIDIA (2024b) NVIDIA. Omniverse, 2024b. URL https://www.nvidia.com/en-us/omniverse/.
- NVIDIA (2024c) NVIDIA. Physx, 2024c. URL https://github.com/NVIDIA-Omniverse/PhysX.
- NVIDIA (2024d) NVIDIA. Edify 3d: Scalable high-quality 3d asset generation. arXiv preprint arXiv:2411.07135, 2024d.
- NVIDIA (2024e) NVIDIA. Transformer engine, 2024e. URL https://github.com/NVIDIA/TransformerEngine.
- OpenAI (2022) OpenAI. Tiktoken, 2022. URL https://github.com/openai/tiktoken.
- OpenAI (2024a) OpenAI. Dall·e 3, 2024a. URL https://openai.com/dall-e. Accessed: [Insert access date here].
- OpenAI (2024b) OpenAI. Sora, 2024b. URL https://openai.com/sora/.
- Pan et al. (2025) Linfei Pan, Dániel Baráth, Marc Pollefeys, and Johannes L Schönberger. Global structure-from-motion revisited. In ECCV, 2025.
- Paszke et al. (2019) Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32, 2019.
- Peebles and Xie (2023) William Peebles and Saining Xie. Scalable diffusion models with transformers. In ICCV, 2023.
- Peng and Quesnelle (2023) Bowen Peng and Jeffrey Quesnelle. Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degradation, 2023.
- Peng et al. (2023) Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. Yarn: Efficient context window extension of large language models. arXiv preprint arXiv:2309.00071, 2023.
- Perazzi et al. (2016) Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In CVPR, 2016.
- Podell et al. (2024) Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. SDXL: Improving latent diffusion models for high-resolution image synthesis. In ICLR, 2024.
- Polyak et al. (2024) Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, et al. Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720, 2024.
- Poole et al. (2023) Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In ICLR, 2023.
- Prasad et al. (2024) Aaditya Prasad, Kevin Lin, Jimmy Wu, Linqi Zhou, and Jeannette Bohg. Consistency policy: Accelerated visuomotor policies via consistency distillation. arXiv preprint arXiv:2405.07503, 2024.
- Qian et al. (2024) Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors. In ICLR, 2024.
- Raffel et al. (2020) Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 2020.
- Ramachandran et al. (2017) Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017.
- Ramesh et al. (2021) Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021.
- Ramesh et al. (2022) Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
- RAPIDS (2023) RAPIDS. Rapids: Libraries for end to end gpu data science, 2023. URL https://rapids.ai.
- Ren et al. (2024) Weiming Ren, Huan Yang, Ge Zhang, Cong Wei, Xinrun Du, Wenhao Huang, and Wenhu Chen. Consisti2v: Enhancing visual consistency for image-to-video generation. arXiv preprint arXiv:2402.04324, 2024.
- Robine et al. (2023) Jan Robine, Marc Höftmann, Tobias Uelwer, and Stefan Harmeling. Transformer-based world models are happy with 100k interactions. arXiv preprint arXiv:2303.07109, 2023.
- Rombach et al. (2021) Robin Rombach, Patrick Esser, and Björn Ommer. Geometry-free view synthesis: Transformers and no 3d priors. In ICCV, 2021.
- Rombach et al. (2022) Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022.
- Runway (2024) Runway. Gen 3, 2024. URL https://runwayml.com/research/introducing-gen-3-alpha.
- Sadat et al. (2024) Seyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges, and Romann M Weber. Litevae: Lightweight and efficient variational autoencoders for latent diffusion models. arXiv preprint arXiv:2405.14477, 2024.
- Saharia et al. (2022) Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022.
- Sajjadi et al. (2022) Mehdi SM Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani Vora, Mario Lučić, Daniel Duckworth, Alexey Dosovitskiy, et al. Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. In CVPR, 2022.
- Sampson (1982) Paul D Sampson. Fitting conic sections to “very scattered” data: An iterative refinement of the bookstein algorithm. Computer graphics and image processing, 1982.
- Schönberger et al. (2016) Johannes L Schönberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys. Pixelwise view selection for unstructured multi-view stereo. In ECCV, 2016.
- Schönberger and Frahm (2016) Johannes Lutz Schönberger and Jan-Michael Frahm. Structure-from-motion revisited. In CVPR, 2016.
- Schuhmann (2022) Christoph Schuhmann. Improved Aesthetic Predictor, 2022. URL https://github.com/christophschuhmann/improved-aesthetic-predictor.
- Shi et al. (2023) Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023.
- Shoeybi et al. (2019) Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
- Simonyan and Zisserman (2014) Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
- Sitzmann et al. (2021) Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neural scene representations with single-evaluation rendering. In NeurIPS, 2021.
- Song et al. (2020) Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
- Soucek and Lokoc (2024) Tomás Soucek and Jakub Lokoc. Transnet v2: An effective deep network architecture for fast shot transition detection. In ACM MM, 2024.
- Su et al. (2024) Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 2024.
- Sun et al. (2024a) Peize Sun, Yi Jiang, Shoufa Chen, Shilong Zhang, Bingyue Peng, Ping Luo, and Zehuan Yuan. Autoregressive model beats diffusion: Llama for scalable image generation. arXiv preprint arXiv:2406.06525, 2024a.
- Sun et al. (2024b) Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang. Generative multimodal models are in-context learners. In CVPR, 2024b.
- Tancik et al. (2023) Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. In ACM SIGGRAPH, 2023.
- Tang et al. (2018) Shitao Tang, Litong Feng, Zhanghui Kuang, Yimin Chen, and Wei Zhang. Fast video shot transition localization with deep structured models. In ACCV, 2018.
- Tatarchenko et al. (2016) Maxim Tatarchenko, Alexey Dosovitskiy, and Thomas Brox. Multi-view 3d models from single images with a convolutional network. In ECCV, 2016.
- Team (2024a) Chameleon Team. Chameleon: Mixed-modal early-fusion foundation models. URL https://arxiv. org/abs/2405.09818, 2024a.
- Team (2024b) Gemma Team. Gemma 2: Improving open language models at a practical size, 2024b. URL https://arxiv.org/abs/2408.00118.
- Technologies (2024) 1X Technologies. 1xgpt, 2024. URL https://github.com/1x-technologies/1xgpt.
- Teed and Deng (2020) Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV, 2020.
- Teed and Deng (2021) Zachary Teed and Jia Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. In NeurIPS, 2021.
- Teng et al. (2024) Yao Teng, Han Shi, Xian Liu, Xuefei Ning, Guohao Dai, Yu Wang, Zhenguo Li, and Xihui Liu. Accelerating auto-regressive text-to-image generation with training-free speculative jacobi decoding. arXiv preprint arXiv:2410.01699, 2024.
- Tucker and Snavely (2020) Richard Tucker and Noah Snavely. Single-view view synthesis with multiplane images. In CVPR, 2020.
- Tulyakov et al. (2018) Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. MoCoGAN: Decomposing motion and content for video generation. In CVPR, 2018.
- Unterthiner et al. (2019) Thomas Unterthiner, Sjoerd van Steenkiste, Karol Kurach, Raphaël Marinier, Marcin Michalski, and Sylvain Gelly. Fvd: A new metric for video generation. In ICLR Workshops, 2019.
- Valevski et al. (2024) Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter. Diffusion models are real-time game engines. arXiv preprint arXiv:2408.14837, 2024.
- van den Oord et al. (2017) Aaron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural discrete representation learning. In NeurIPS, 2017.
- Vaswani et al. (2017) Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
- Villegas et al. (2023) Ruben Villegas, Mohammad Babaeizadeh, Pieter-Jan Kindermans, Hernan Moraldo, Han Zhang, Mohammad Taghi Saffar, Santiago Castro, Julius Kunze, and Dumitru Erhan. Phenaki: Variable length video generation from open domain textual description. In ICLR, 2023.
- Walke et al. (2023) Homer Walke, Kevin Black, Abraham Lee, Moo Jin Kim, Max Du, Chongyi Zheng, Tony Zhao, Philippe Hansen-Estruch, Quan Vuong, Andre He, Vivek Myers, Kuan Fang, Chelsea Finn, and Sergey Levine. Bridgedata v2: A dataset for robot learning at scale. In CoRL, 2023.
- Wang et al. (2024a) Junke Wang, Yi Jiang, Zehuan Yuan, Binyue Peng, Zuxuan Wu, and Yu-Gang Jiang. Omnitokenizer: A joint image-video tokenizer for visual generation. arXiv preprint arXiv:2406.09399, 2024a.
- Wang et al. (2021a) Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. In NeurIPS, 2021a.
- Wang et al. (2024b) Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024b.
- Wang et al. (2018) Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. Video-to-video synthesis. In NeurIPS, 2018.
- Wang et al. (2019) Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Jan Kautz, and Bryan Catanzaro. Few-shot video-to-video synthesis. In NeurIPS, 2019.
- Wang et al. (2021b) Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. One-shot free-view neural talking-head synthesis for video conferencing. In CVPR, 2021b.
- Wang et al. (2024c) Xiang Wang, Shiwei Zhang, Changxin Gao, Jiayu Wang, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, and Nong Sang. Unianimate: Taming unified video diffusion models for consistent human image animation. arXiv preprint arXiv:2406.01188, 2024c.
- Wang et al. (2023a) Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, Jiagang Zhu, and Jiwen Lu. Drivedreamer: Towards real-world-driven world models for autonomous driving. arXiv preprint arXiv:2309.09777, 2023a.
- Wang et al. (2024d) Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, et al. Emu3: Next-token prediction is all you need. arXiv preprint arXiv:2409.18869, 2024d.
- Wang et al. (2023b) Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video generation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023b.
- Wang et al. (2025) Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, et al. Internvideo2: Scaling foundation models for multimodal video understanding. In ECCV, 2025.
- Wang et al. (2024e) Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, and Zhaoxiang Zhang. Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving. In CVPR, 2024e.
- Wang et al. (2024f) Zhendong Wang, Zhaoshuo Li, Ajay Mandlekar, Zhenjia Xu, Jiaojiao Fan, Yashraj Narang, Linxi Fan, Yuke Zhu, Yogesh Balaji, Mingyuan Zhou, et al. One-step diffusion policy: Fast visuomotor policies via diffusion distillation. arXiv preprint arXiv:2410.21257, 2024f.
- Wang et al. (2024g) Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH, 2024g.
- Weng et al. (2023) Qizhen Weng, Lingyun Yang, Yinghao Yu, Wei Wang, Xiaochuan Tang, Guodong Yang, and Liping Zhang. Beware of fragmentation: Scheduling $\{$GPU-Sharing$\}$ workloads with fragmentation gradient descent. In USENIX ATC, 2023.
- Wiles et al. (2020) Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. Synsin: End-to-end view synthesis from a single image. In CVPR, 2020.
- Wortsman et al. (2023) Mitchell Wortsman, Peter J Liu, Lechao Xiao, Katie Everett, Alex Alemi, Ben Adlam, John D Co-Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, et al. Small-scale proxies for large-scale transformer training instabilities. arXiv preprint arXiv:2309.14322, 2023.
- Wu et al. (2022) Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. Nüwa: Visual synthesis pre-training for neural visual world creation. In ECCV, 2022.
- Wu et al. (2023a) Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, and Weisi Lin. Exploring video quality assessment on user generated contents from aesthetic and technical perspectives. In ICCV, 2023a.
- Wu et al. (2023b) Philipp Wu, Alejandro Escontrela, Danijar Hafner, Pieter Abbeel, and Ken Goldberg. Daydreamer: World models for physical robot learning. In CoRL, 2023b.
- Wu et al. (2024) Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Haotian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, et al. Vila-u: a unified foundation model integrating visual understanding and generation. arXiv preprint arXiv:2409.04429, 2024.
- Wu and He (2018) Yuxin Wu and Kaiming He. Group normalization. In ECCV, 2018.
- Xu et al. (2024) Dejia Xu, Weili Nie, Chao Liu, Sifei Liu, Jan Kautz, Zhangyang Wang, and Arash Vahdat. Camco: Camera-controllable 3d-consistent image-to-video generation. arXiv preprint arXiv:2406.02509, 2024.
- Xu et al. (2019) Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin. Understanding and improving layer normalization. In NeurIPS, 2019.
- Xue et al. (2024) Fuzhao Xue, Yukang Chen, Dacheng Li, Qinghao Hu, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Zhijian Liu, et al. Longvila: Scaling long-context visual language models for long videos. arXiv preprint arXiv:2408.10188, 2024.
- Yan et al. (2021) Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:2104.10157, 2021.
- Yang et al. (2024a) An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115, 2024a.
- Yang et al. (2024b) Cheng-Yen Yang, Hsiang-Wei Huang, Wenhao Chai, Zhongyu Jiang, and Jenq-Neng Hwang. Samurai: Adapting segment anything model for zero-shot visual tracking with motion-aware memory. arXiv preprint arXiv:2411.11922, 2024b.
- Yang et al. (2024c) Jiazhi Yang, Shenyuan Gao, Yihang Qiu, Li Chen, Tianyu Li, Bo Dai, Kashyap Chitta, Penghao Wu, Jia Zeng, Ping Luo, et al. Generalized predictive model for autonomous driving. In CVPR, 2024c.
- Yang et al. (2023) Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel. Learning interactive real-world simulators. arXiv preprint arXiv:2310.06114, 2023.
- Yang et al. (2024d) Zhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding, Shiyu Huang, Jiazheng Xu, Yuanming Yang, Wenyi Hong, Xiaohan Zhang, Guanyu Feng, et al. Cogvideox: Text-to-video diffusion models with an expert transformer. arXiv preprint arXiv:2408.06072, 2024d.
- Yin et al. (2024) Tianwei Yin, Qiang Zhang, Richard Zhang, William T Freeman, Fredo Durand, Eli Shechtman, and Xun Huang. From slow bidirectional to fast causal video generators. arXiv preprint arXiv:2412.07772, 2024.
- Yu et al. (2021) Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021.
- Yu et al. (2020) Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In CVPR, 2020.
- Yu et al. (2022) Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. TMLR, 2022.
- Yu et al. (2023a) Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, and Lu Jiang. MAGVIT: Masked generative video transformer. In CVPR, 2023a.
- Yu et al. (2024a) Lijun Yu, Jose Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A Ross, and Lu Jiang. Language model beats diffusion - tokenizer is key to visual generation. In ICLR, 2024a.
- Yu et al. (2024b) Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. arXiv preprint arXiv:2406.07550, 2024b.
- Yu et al. (2023b) Sihyun Yu, Kihyuk Sohn, Subin Kim, and Jinwoo Shin. Video probabilistic diffusion models in projected latent space. In CVPR, 2023b.
- Zeng et al. (2024) Yan Zeng, Guoqiang Wei, Jiani Zheng, Jiaxin Zou, Yang Wei, Yuchen Zhang, and Hang Li. Make pixels dance: High-dynamic video generation. In CVPR, 2024.
- Zhai et al. (2023) Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In ICCV, 2023.
- Zhang and Sennrich (2019) Biao Zhang and Rico Sennrich. Root mean square layer normalization. In NeurIPS, 2019.
- Zhang et al. (2018) Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018.
- Zhang et al. (2024) Weipu Zhang, Gang Wang, Jian Sun, Yetian Yuan, and Gao Huang. Storm: Efficient stochastic transformer based world models for reinforcement learning. In NeurIPS, 2024.
- Zhao et al. (2024a) Guosheng Zhao, Xiaofeng Wang, Zheng Zhu, Xinze Chen, Guan Huang, Xiaoyi Bao, and Xingang Wang. Drivedreamer-2: Llm-enhanced world models for diverse driving video generation. arXiv preprint arXiv:2403.06845, 2024a.
- Zhao et al. (2024b) Yue Zhao, Yuanjun Xiong, and Philipp Krähenbühl. Image and video tokenization with binary spherical quantization. arXiv preprint arXiv:2406.07548, 2024b.
- Zhou et al. (2024a) Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, and Omer Levy. Transfusion: Predict the next token and diffuse images with one multi-modal model. arXiv preprint arXiv:2408.11039, 2024a.
- Zhou et al. (2024b) Siyuan Zhou, Yilun Du, Jiaben Chen, YANDONG LI, Dit-Yan Yeung, and Chuang Gan. Robodreamer: Learning compositional world models for robot imagination. In ICML, 2024b.
- Zhou et al. (2018) Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: learning view synthesis using multiplane images. ACM Transactions on Graphics (TOG), 2018.
- Zhu et al. (2024) Fangqi Zhu, Hongtao Wu, Song Guo, Yuxiao Liu, Chilam Cheang, and Tao Kong. Irasim: Learning interactive real-robot action simulators. arXiv preprint arXiv:2406.14540, 2024.
- Zhu et al. (2023) Wentao Zhu, Yufang Huang, Xiufeng Xie, Wenxian Liu, Jincan Deng, Debing Zhang, Zhangyang Wang, and Ji Liu. Autoshot: A short video dataset and state-of-the-art shot boundary detection. In CVPR Workshops, 2023.
術語對照表
| 英文 | 中文 |
|---|---|
| Physical AI | 物理 AI |
| World Model | 世界模型 |
| World Foundation Model (WFM) | 世界基礎模型 |
| Digital Twin | 數位孿生 |
| Policy Model | 策略模型 |
| Policy Evaluation | 策略評估 |
| Policy Initialization | 策略初始化 |
| Policy Training | 策略訓練 |
| Model-Predictive Control | 模型預測控制 |
| Receding Horizon | 滾動時域 |
| Synthetic Data Generation | 合成資料生成 |
| Reward Model | 獎勵模型 |
| Reinforcement Learning | 強化學習 |
| Agent | 代理人 |
| Sensor | 感測器 |
| Actuator | 致動器 |
| Perturbation | 擾動 |
| Video Curation Pipeline | 影片整理流程 |
| Data Curation | 資料整理 |
| Shot Detection | 鏡頭偵測 |
| Shot Boundary Detection | 鏡頭邊界偵測 |
| Shot Transition | 鏡頭轉換 |
| Transcoding | 轉檔 |
| Filtering | 過濾 |
| Annotation | 標註 |
| Deduplication | 去重 |
| Semantic Deduplication | 語意去重 |
| Sharding | 分片 |
| Motion Vector | 運動向量 |
| Optical Flow | 光流 |
| Taxonomy | 分類法 |
| Alt Text | 替代文字 |
| Embedding | 嵌入 |
| Tokenization | Token 化 |
| Tokenizer | 標記器(本文保留原文) |
| Latent Space | 潛在空間 |
| Latent Representation | 潛在表徵 |
| Latent Code | 潛在碼 |
| Latent Diffusion Model | 潛在擴散模型 |
| Quantized Index | 量化索引 |
| Quantizer | 量化器 |
| Vocabulary | 詞彙表 |
| Autoencoder (AE) | 自編碼器 |
| Variational Autoencoder (VAE) | 變分自編碼器 |
| Finite-Scalar-Quantization (FSQ) | 有限純量量化 |
| Wavelet Transform | 小波轉換 |
| Wavelet Space | 小波空間 |
| Residual Block | 殘差區塊 |
| Causal Temporal Convolution | 因果時間卷積 |
| Causal Self-attention | 因果自注意力 |
| Left Padding | 左側填補 |
| Layer Normalization (LayerNorm) | 層正規化 |
| Group Normalization (GroupNorm) | 群組正規化 |
| Root Mean Square Normalization (RMSNorm) | 均方根正規化 |
| Perceptual Loss | 感知損失 |
| Adversarial Loss | 對抗損失 |
| Commitment Loss | 承諾損失 |
| Cross-entropy Loss | 交叉熵損失 |
| Negative Log-likelihood (NLL) | 負對數概似 |
| Denoising | 去噪 |
| Denoiser | 去噪器 |
| Denoising Score Matching | 去噪分數匹配 |
| Diffusion Model | 擴散模型 |
| Diffusion Decoder | 擴散解碼器 |
| Autoregressive Model | 自迴歸模型 |
| Gaussian Flow Matching | 高斯流匹配 |
| Preconditioning | 前置條件化 |
| Multi-task Learning | 多任務學習 |
| Patchification | Patch 化 |
| Rotary Position Embedding (RoPE) | 旋轉位置嵌入 |
| Absolute Positional Embedding (APE) | 絕對位置嵌入 |
| Adaptive Layer Normalization (AdaLN) | 自適應層正規化 |
| Low-Rank Adaptation (LoRA) | 低秩適應 |
| Cross-attention | 交叉注意力 |
| Attention Mechanism | 注意力機制 |
| Attention Entropy Collapse | 注意力熵崩塌 |
| Classifier-free Guidance | 無分類器引導 |
| Negative Prompt | 負向提示詞 |
| Prompt | 提示詞 |
| Prompt Upsampler | 提示詞上採樣器 |
| Zero-shot Prompt Engineering | 零樣本提示工程 |
| Pre-training | 預訓練 |
| Post-training | 後訓練 |
| Fine-tuning | 微調 |
| Progressive Training | 漸進式訓練 |
| Cooling-down | 冷卻階段 |
| Training Curriculum | 訓練課程 |
| Mini-batch | 小批次 |
| Bucket | 桶 |
| Reflection Padding | 鏡射填補 |
| Longest-side Resizing | 最長邊縮放 |
| Loss Spike | 損失尖峰 |
| Exponential Moving Average (EMA) | 指數移動平均 |
| Activation Checkpointing | 激活檢查點 |
| Fully Sharded Data Parallelism (FSDP) | 全分片資料平行 |
| Context Parallelism (CP) | 上下文平行 |
| Tensor Parallelism (TP) | 張量平行 |
| Sequence Parallelism (SP) | 序列平行 |
| Model FLOPs Utilization (MFU) | 模型 FLOPs 利用率 |
| Speculative Decoding | 推測式解碼 |
| Decoding Head | 解碼頭 |
| Rejection Sampling | 拒絕取樣 |
| Unembedding Layer | 反嵌入層 |
| Catastrophic Forgetting | 災難性遺忘 |
| Byte Pair Encoding (BPE) | 位元組對編碼 |
| Foresight Generation | 前瞻生成 |
| 3D Consistency | 3D 一致性 |
| Physics Alignment | 物理對齊 |
| Intuitive Physics | 直覺物理 |
| Rigid Body Dynamics | 剛體動力學 |
| Epipolar Geometry | 對極幾何 |
| Fundamental Matrix | 基礎矩陣 |
| Multi-view Geometry | 多視角幾何 |
| Novel Viewpoint | 新視角 |
| 3D Gaussian Splatting | 3D 高斯潑濺 |
| Structure-from-Motion | 運動恢復結構 |
| Instance Segmentation Mask | 實例分割遮罩 |
| Intersection-over-Union (IoU) | 交並比 |
| Ground Truth | 真值 |
| Camera Pose | 相機姿態 |
| View Embedding | 視角嵌入 |
| Ego-motion | 自車運動 |
| Multi-view | 多視角 |
| Trajectory | 軌跡 |
| Absolute Trajectory Error (ATE) | 絕對軌跡誤差 |
| Temporal Sampson Error (TSE) | 時間 Sampson 誤差 |
| Cross-view Sampson Error (CSE) | 跨視角 Sampson 誤差 |
| Trajectory Agreement Error (TAE) | 軌跡吻合誤差 |
| Trajectory Following Error (TFE) | 軌跡遵循誤差 |
| Dense Bundle Adjustment | 稠密光束法平差 |
| Undistort | 去畸變 |
| End-effector | 末端執行器 |
| Visuomotor Policy | 視覺運動策略 |
| Imitation Learning | 模仿學習 |
| Embodied AI | 具身 AI |
| Humanoid | 人形機器人 |
| Action Embedder | 動作嵌入器 |
| Timestamp Embedding | 時間戳嵌入 |
| Guardrail | 護欄 |
| Red Team | 紅隊 |
| Lemmatization | 詞形還原 |
| Pixelation | 像素化 |
| Corner Case | 極端案例 |
| Adversarial | 對抗性 |
| Object Permanence | 物體恆存性 |
| Contact-rich Dynamics | 接觸密集動態 |
| Sim-to-Real / Sim2Real | 模擬到真實 |
| Domain Gap | 領域差距 |
| Distribution Shift | 分布偏移 |
| Out-of-distribution | 分布外 |
| Generalist | 通才 |
| Generative Simulation | 生成式模擬 |
| Neural Rendering | 神經渲染 |
| Recurrent Neural Network | 循環神經網路 |
| System Identification | 系統辨識 |
| Video Codec | 影片編解碼器 |
| Dataloader | 資料載入器 |
| Audio Remixing | 音訊重混 |
| Upsampling | 上採樣 |
| Downsampling | 下採樣 |
| Sufficient Statistics | 充分統計量 |
| Isotropic Gaussian Prior | 等向性高斯先驗分布 |
| Non-stationary | 非平穩 |
| Scale Invariance | 尺度不變性 |
| Morphing Artifact | 形變瑕疵 |
| Artifact | 瑕疵 |
| Distortion | 失真 |
| Peak Signal-to-Noise Ratio (PSNR) | 峰值訊噪比 |
| Structural Similarity (SSIM) | 結構相似度 |
| Mode-seeking | 尋峰 |
| Augmented Noise | 增強雜訊 |
| Pan / Zoom / Tilt | 平移/變焦/俯仰 |
































































































































































