The Rise of Verbal Reinforcement Learning
2026-09-01 • Computation and Language
Computation and LanguageArtificial Intelligence
AI summaryⓘ
The authors explain a new way of teaching language-based AI agents using natural human language as feedback, which they call Verbal Reinforcement Learning (VRL). They organize this approach into three main types: using language to set goals and rewards, to guide decision-making during testing, and to help train the model itself. Their framework helps understand how language influences an agent’s behavior at different stages. They also discuss challenges and future directions for improving these language-guided agents.
Natural Language ProcessingReinforcement LearningLanguage ModelsFeedback SignalAgent TrainingGoal SpecificationDeliberative FeedbackModel ParametersTask Grounding
Authors
Kshitij Tayal, Arun Sharma, Genta Indra Winata, Anirban Das, Sambit Sahu
Abstract
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and \textit{what} it modifies, yielding three pillars: (1) \textbf{Language as Grounding Signal}, where language defines the task itself by specifying goals, states, and reward structures; (2) \textbf{Language as Deliberative Feedback}, where natural language guides reasoning at test time without the need to update model parameters; (3) \textbf{Language as Learning Signal}, where language-based feedback shapes model parameters through training. Within each pillar, we synthesize representative work, distinguish key subcategories of approaches, and outline the distinct role language plays in shaping agent behavior. Together, this taxonomy shows how verbal reinforcement is reshaping agent development, while also defining the challenges and opportunities for building more capable and aligned agents.