MerchantBench Reveals Massive Gap in Long-Term Agent Coherence
A new benchmark from top Chinese universities, MerchantBench, simulates 365-day e-commerce tasks to test LLM agents' long-term coherence. The best model (GPT-5.6 Sol) achieves only 27.3% of human performance, highlighting a fundamental limitation in sustained decision-making for real-world deployment.
Sources (2)
Updated Aug 7, 2026