← Back to brief
ResearchOfficialPreprintarXiv AI/ML

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

A new benchmark, MerchantBench, simulates a full year of e-commerce operations to evaluate large language model (LLM) agents' ability to maintain coherent, purposeful behavior over long time horizons. Using real product data and a suite of 26 tools, the benchmark tests LLMs on complex, interdependent tasks such as sourcing, pricing, and cash-flow management. Results show that the best-performing LLM agent achieves only 27.3% of the average final net assets of human participants, indicating a significant performance gap.

Why it matters: This work highlights a major limitation of current LLM agents in handling long-term, real-world decision-making tasks, which is critical for their practical deployment in complex domains like e-commerce.

Full story at: arXiv AI/ML