ARC-AGI-3 A New Challenge for Frontier Agentic Intelligence
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
Thinking Augmented Pre-training
Co-Evolving Policy Distillation
OpenThoughts Data Recipes for Reasoning Models
Self-Distilled RLVR
Hybrid Policy Distillation for LLMs
How to Fine-Tune a Reasoning Model A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data
Self-Anchor LLM Reasoning via Step-by-step Attention Alignment
Lightning OPD Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
Interleaved Head Attention
Rethinking Generalization in Reasoning SFT
ReAct Synergizing Reasoning and Acting in Language Models