MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
2026-08-03 • Artificial Intelligence
Artificial IntelligenceComputers and Society
AI summaryⓘ
The authors created MonitrLLM, a tool that connects what users say to language models with their personal goals and success ratings. They tested it with college students using ChatGPT and found that even when people felt satisfied, their tasks still failed about 23% of the time. They also noticed that longer conversations often meant more difficulty instead of better engagement. The study suggests tracking both the full chat and user feedback helps better understand how well language models perform.
LLM evaluationuser feedbackconversation transcriptstask outcomesmulti-turn conversationsChatGPTinteraction trajectoriesuser satisfactionbenchmark suitesnaturalistic use
Authors
Victor Ojewale, Ro Encarnación, Suresh Venkatasubramanian, Danaé Metaxa
Abstract
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes. For instance, despite reporting high average satisfaction (4.19/5) with their LLM interactions, participants also experience a substantial 23.1% failure rate on their goal tasks. We also find that multi-turn conversations are reported as failing at 2.5 times the rate of single-turn exchanges, a pattern that reframes extended interaction as a signal of difficulty rather than engagement. We conclude by discussing the value of incorporating direct user feedback with observational data for robust LLM evaluations, and the possibilities for infrastructure that enables this goal.