Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Muon is Scalable for LLM Training
Kimi K3: Open Frontier Intelligence
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Continual Pre-training of MoEs: How robust is your router?
In-Place Tokenizer Expansion for Pre-trained LLMs
Instruction-Following Pruning for Large Language Models
Nemotron 3 Nano Open Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Nemotron 3 Super Open Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
Qwen3 Technical Report