Invent a Dataset: Measuring dataset generation abilities with zero seed
Harness-Zero: Harness Distillation via Agent-as-Harness
Scaling Data-Constrained Language Models
Muon is Scalable for LLM Training
Kimi K3: Open Frontier Intelligence
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Joint Optimization of Tool Creation and Use for Large Language Model Agents
Continual Pre-training of MoEs: How robust is your router?
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
In-Place Tokenizer Expansion for Pre-trained LLMs
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Instruction-Following Pruning for Large Language Models
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use
EndPrompt: Efficient Long-Context Extension via Terminal Anchoring
MeMo: Memory as a Model
Small LLMs: Pruning vs. Training from Scratch
TW-LegalBench: Measuring Taiwanese Legal Understanding
Thinking Augmented Pre-training
Co-Evolving Policy Distillation
Language Models are Few-Shot Learners
DataComp-LM In search of the next generation of training sets for language models
OpenThoughts Data Recipes for Reasoning Models
Self-Distilled RLVR
Attention to Mamba A Recipe for Cross-Architecture Distillation
Autogenesis A Self-Evolving Agent Protocol
LLMs Corrupt Your Documents When You Delegate
Hybrid Policy Distillation for LLMs
Turning the TIDE Cross-Architecture Distillation for Diffusion Large Language Models
Prefill-as-a-Service KVCache of Next-Generation Models Could Go Cross-Datacenter
Training LLM Agents for Spontaneous Reward-Free Self-Evolution via World Knowledge Exploration
Nemotron 3 Nano Open Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
How to Fine-Tune a Reasoning Model A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
LongAct Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
TREX Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
DFlash Block Diffusion for Flash Speculative Decoding
Self-Anchor LLM Reasoning via Step-by-step Attention Alignment
Context Parallelism for Scalable Million-Token Inference
SERA Soft-Verified Efficient Repository Agents
Self-Distillation Enables Continual Learning
Lightning OPD Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Nemotron 3 Super Open Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Large Language Model Post-Training A Unified View of Off-Policy and On-Policy Learning
Learning is Forgetting LLM Training As Lossy Compression
DoReMi Optimizing Data Mixtures Speeds Up Language Model Pretraining
Interleaved Head Attention
Rethinking Generalization in Reasoning SFT
Scaling Latent Reasoning via Looped Language Models
Functionality-Oriented LLM Merging on the Fisher-Rao Manifold
ReAct Synergizing Reasoning and Acting in Language Models
Better and Faster Large Language Models via Multi-token Prediction
In-Place Test-Time Training
Automating Database-Native Function Code Synthesis with LLMs
2 OLMo 2 Furious
Self-Improving Pretraining using post-trained models to pretrain better models
Qwen3 Technical Report