AI Evaluation Is Becoming a Procurement and Reliability Battleground
Vals, backed by Andreessen Horowitz, is positioning confidential, domain-specific benchmarks and negative-outcome measurement as an alternative to easily gamed public tests. Recent GPT-5.5 versus Qwen3.7 Max comparisons reinforce that model leadership is workload-specific, with overlapping intervals and provisional or non-comparable metrics. Independent validation, reproducibility, and benchmark security are becoming increasingly important for buyers and builders.
Sources (2)
Updated Sep 19, 2026