Agentic systems improve reasoning fairness and efficiency
Agentic Multi-Turn Reasoning: A Fairness Approach
Artificial IntelligenceMachine Learning
Summary
Solving complex problems often requires reasoning through many steps, but teaching AI to do this well is tricky when feedback only comes after all steps are done. It’s even harder when the data mainly shows common reasoning paths, making rare but important approaches overlooked. The authors developed a new method called Fair Multi-Level Preference Optimization that helps AI models learn better by handling this long-step challenge and data imbalance fairly. Their approach showed improved performance on benchmarks designed to test step-by-step AI reasoning.
What this means in practice
- •For ai developers: Improve multi-step reasoning agents by reducing bias against rare reasoning paths, enhancing accuracy on complex tasks.
- •For software engineers: Build more reliable AI assistants that plan and verify multi-step actions effectively despite imbalanced training data.
- •For automated customer service teams: Deploy chatbots capable of better step-by-step problem solving using fair preference optimization to handle diverse query types.$Commercial implications: Enables development of commercially viable AI chatbots with improved reasoning for customer support applications.
Authors
Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren, Marios Savvides, Khoa Luu
Abstract
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $Φ$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.