RLAES: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
Researchers introduce RLAES, a unified large language model framework that jointly optimizes automated essay scoring and feedback generation using reinforcement learning. The system incorporates Rubric-based Feedback Evaluation (RFE), which uses 166 fine-grained binary rubric items and an LLM-as-judge to make feedback quality measurable. On the ASAP benchmark, RLAES achieves the highest scoring performance among LLM-based methods (QWK = 0.803) while maintaining feedback quality comparable to GPT-5.5.
Why it matters: This work demonstrates a significant advance in automated essay scoring and feedback generation by leveraging reinforcement learning and rubric-based evaluation to achieve state-of-the-art performance without sacrificing feedback quality.
Full story at: arXiv Computation and Language ↗