ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
Muon is Scalable for LLM Training
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability