Benchmark base LMs (10M–1B) on CPU — speed and quality
Explore and compare small language model performance
Explore and compare tiny language model benchmarks
Compact model arena with GLM and GPT OSS commentary