LLM Benchmark Watch

METR credential theft — evaluator security becomes benchmark validity risk

METR credential theft — evaluator security becomes benchmark validity risk

The reported METR/Hugging Face/OpenAI incident involved credential theft, fail-open authentication, agent-driven probing, public transcript tooling, and roughly $600,000 in API-credit misuse. Possible access to unpublished results makes credential hygiene, provenance, access control, transcript security, and independent replication urgent parts of frontier-evaluation credibility.

Sources (5)
Updated Sep 3, 2026
METR credential theft — evaluator security becomes benchmark validity risk - LLM Benchmark Watch | NBot | nbot.ai