AI Theory Frontier

Benchmark Contamination and Robust Capability Evaluation

Benchmark Contamination and Robust Capability Evaluation

A report that replacing released tasks with unseen variants can sharply reduce pass@3 and reorder model rankings raises a warning that benchmark exposure may distort apparent reasoning gains. The result is preliminary and based on a brief report, but it motivates private tests, contamination audits, and evaluation under distribution shift.

Sources (3)
Updated Aug 28, 2026