Language model agents design their own evaluators to improve task success
Self-Designed Evaluators and Warm Memory for Long-Horizon Agents
Artificial Intelligence
Summary
When a language model agent works on many tasks in a row without clear success signals, it can't tell if it did well or try again safely. The researchers created SelfSuite, where the agent uses its own model to build a set of small tests and judges to check its work. This system helps the agent decide when to retry tasks and learn from past results without needing labels from experts. SelfSuite performs better than agents without such feedback and is competitive with methods that rely on expert labels.
What this means in practice
- •For software developers: Improve autonomous language-based systems by enabling them to self-evaluate and retry tasks without expert feedback.
- •For automated customer service teams: Use self-designed evaluators to enhance chatbot performance over long interactions without needing extensive manual supervision.
Authors
Saeid Asgari, Emre Kiciman, Leonardo de Oliveira Nunes, Ranveer Chandra
Abstract
A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materials, designs a small evaluation suite of weighted judges and grounded per-task briefs, freezes it, and uses it to gate a keep-best retry and to label a typed, outcome-tracked memory. On matched five-repeat benchmarks over tau2-bench and AppWorld, SelfSuite scores above the plain agent without any labels, matches methods given ten expert labels on tau2-bench, and trails Agentic Context Engineering (ACE) on AppWorld, where code execution gives a direct success signal. In an ablation campaign run on the same tasks, it is above label-free ACE in every repeat, and the gated second attempt is the only component whose removal hurts in every repeat. We also simulate a subject-matter expert who grades ten onboarding tasks per world. Using those labels to calibrate SelfSuite's evaluator gives a small, consistent gain, and using them to warm up ACE's memory lifts ACE to tie calibrated SelfSuite. A single-run study on a second model family shows the same ordering.