Trust guided transformer improves long decision making in AI models
Trust Guided Decision Transformer
Machine Learning
Summary
When AI models try to make many decisions in a row, their performance often drops because they rely on past information that becomes less reliable over time. The authors found that the model’s own errors in predicting the next state reveal when this happens. They made a method called Trust Guided Decision Transformer that ignores unreliable past information and uses trustworthy parts to decide the best action. This method improves performance on navigation and movement tasks compared to older approaches.
What this means in practice
- •For robotics engineers: Improve robot control by selecting only reliable past data for better long-term decision making.
- •For autonomous vehicle developers: Enhance autonomous navigation systems by preventing error accumulation using trust-based context selection.
Authors
Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath
Abstract
Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.