System One benchmark brings reproducible evaluation to specialized decision models
A new evaluation covers 346,009 requests across 37 datasets, comparing System One with Qwen3.8-27B and Gemma-4-E4B on accuracy, calibration, and selective prediction. The released code, harness, and raw responses make this unusually reproducible, though contamination and independent replication remain open questions.
Sources (2)
Updated Sep 30, 2026