← Back to brief
ResearchOfficialPreprintarXiv AI/ML

Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making

A new LLM agent architecture is proposed that combines POMDP routing with a self-correcting reward model to address challenges in long-horizon planning and dynamic interaction. Experiments on ALFWorld and WebShop benchmarks show a 24.5% absolute improvement in task success rate over the ReAct baseline. The approach integrates reinforcement learning and graph-based memory, with ablation studies confirming the reward-driven critique module's role in reducing hallucination.

Why it matters: This work presents a practical and scalable framework that advances the reliability and effectiveness of autonomous LLM agents in complex, multi-step environments.

Full story at: arXiv AI/ML