FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
Agent Seer: Synthesizing Scenarios from Specification Understanding
ARC-AGI-3 A New Challenge for Frontier Agentic Intelligence
AutoData: A Multi-Agent System for Open Web Data Collection
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
TW-LegalBench: Measuring Taiwanese Legal Understanding
DataComp-LM In search of the next generation of training sets for language models
OpenThoughts Data Recipes for Reasoning Models
LLMs Corrupt Your Documents When You Delegate
MinerU2.5-Pro Pushing the Limits of Data-Centric Document Parsing at Scale