AI Frontier Updates

MerchantBench Reveals Massive Gap in Long-Term Agent Coherence

MerchantBench Reveals Massive Gap in Long-Term Agent Coherence

A new benchmark from top Chinese universities, MerchantBench, simulates 365-day e-commerce tasks to test LLM agents' long-term coherence. The best model (GPT-5.6 Sol) achieves only 27.3% of human performance, highlighting a fundamental limitation in sustained decision-making for real-world deployment.

Sources (2)
Updated Aug 7, 2026