Invent a Dataset: Measuring dataset generation abilities with zero seed
Harness-Zero: Harness Distillation via Agent-as-Harness
A Survey of Agentic Reasoning for Large Language Models
DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale
Reinforcing Agents with Collective Skills
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
FrogNano: Training a 4B Coding Agent via Online Task Synthesis
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
Agent Seer: Synthesizing Scenarios from Specification Understanding
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Cut Your Losses in Large-Vocabulary Language Models
Gated Delta Networks: Improving Mamba2 with Delta Rule
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
An Empirical Model of Large-Batch Training
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
Scaling Data-Constrained Language Models
Muon is Scalable for LLM Training
Kimi K3: Open Frontier Intelligence
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Joint Optimization of Tool Creation and Use for Large Language Model Agents
Continual Pre-training of MoEs: How robust is your router?
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
Cosmos World Foundation Model Platform for Physical AI
Cosmos 3 Omnimodal World Models for Physical AI
ARC-AGI-3 A New Challenge for Frontier Agentic Intelligence
In-Place Tokenizer Expansion for Pre-trained LLMs
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Autodata: An agentic data scientist to create high quality synthetic data
Instruction-Following Pruning for Large Language Models
AutoData: A Multi-Agent System for Open Web Data Collection
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
MeMo: Memory as a Model
Small LLMs: Pruning vs. Training from Scratch
TW-LegalBench: Measuring Taiwanese Legal Understanding
Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context
Thinking Augmented Pre-training
Co-Evolving Policy Distillation
Language Models are Few-Shot Learners
DataComp-LM In search of the next generation of training sets for language models
OpenThoughts Data Recipes for Reasoning Models
Self-Distilled RLVR
Pioneer Agent Continual Improvement of Small Language Models in Production
Attention to Mamba A Recipe for Cross-Architecture Distillation
Autogenesis A Self-Evolving Agent Protocol
LLMs Corrupt Your Documents When You Delegate
Hybrid Policy Distillation for LLMs
Turning the TIDE Cross-Architecture Distillation for Diffusion Large Language Models
Prefill-as-a-Service KVCache of Next-Generation Models Could Go Cross-Datacenter
Training LLM Agents for Spontaneous Reward-Free Self-Evolution via World Knowledge Exploration
Nemotron 3 Nano Open Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
How to Fine-Tune a Reasoning Model A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
Dive into Claude Code The Design Space of Todays and Future AI Agent Systems
LongAct Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
MinerU2.5-Pro Pushing the Limits of Data-Centric Document Parsing at Scale
TREX Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
DFlash Block Diffusion for Flash Speculative Decoding
Self-Anchor LLM Reasoning via Step-by-step Attention Alignment
Context Parallelism for Scalable Million-Token Inference
Improved Alignment of Modalities in Large Vision Language Models
SERA Soft-Verified Efficient Repository Agents
Self-Distillation Enables Continual Learning
Lightning OPD Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Nemotron 3 Super Open Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Large Language Model Post-Training A Unified View of Off-Policy and On-Policy Learning
Learning is Forgetting LLM Training As Lossy Compression
DoReMi Optimizing Data Mixtures Speeds Up Language Model Pretraining
VisionFoundry Teaching VLMs Visual Perception with Synthetic Images
Interleaved Head Attention
From Safety Risk to Design Principle Peer-Preservation in Multi-Agent LLM Systems
Externalization in LLM Agents A Unified Review of Memory Skills Protocols and Harness Engineering
Rethinking Generalization in Reasoning SFT
Challenges and Research Directions for Large Language Model Inference Hardware
Scaling Latent Reasoning via Looped Language Models
Neural Computers
dots.ocr Multilingual Document Layout Parsing in a Single Vision-Language Model
Functionality-Oriented LLM Merging on the Fisher-Rao Manifold
HDP A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems
ReAct Synergizing Reasoning and Acting in Language Models
Better and Faster Large Language Models via Multi-token Prediction
Transformers are SSMs Generalized Models and Efficient Algorithms Through Structured State Space Duality
In-Place Test-Time Training
Automating Database-Native Function Code Synthesis with LLMs
Act Wisely Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
2 OLMo 2 Furious
Olmo 3
Edge Intelligence with Spiking Neural Networks
The Art of Building Verifiers for Computer Use Agents
Self-Improving Pretraining using post-trained models to pretrain better models
The Depth Ceiling On the Limits of Large Language Models in Discovering Latent Planning
TurboQuant Online Vector Quantization with Near-optimal Distortion Rate
Qwen3 Technical Report
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge