Vision language action policies gain reliable refusal option under uncertainty
Do Not Cut When Uncertain: Rejectable and Calibrated Decision Heads for VLA Policies in Robotic Harvesting
Robotics
Summary
Robots that pick fruit can make costly mistakes when they cannot clearly see if a stem is safe to cut. The researchers found that the problem isn’t how big or complex the robot’s brain is, but how it decides whether to act or not. They created a way for the robot to say "I don’t know" and avoid making risky cuts when unsure. Their method works by adding a special decision layer on top of the existing robot system without retraining it. This helps the robot avoid errors in confusing situations while still performing well when it is sure.
What this means in practice
- •For robotics engineers: Enable robotic harvesters to abstain from cutting when uncertain, reducing costly errors caused by occluded stems without retraining the whole model.
- •For industrial automation teams: Integrate a rejectable decision output in existing vision-based control systems to improve safety and reliability under visual uncertainty.
Authors
Heng Zhang
Abstract
Vision-Language-Action (VLA) policies trained with behavior cloning or flow matching are optimized to output an action trajectory, but they cannot express "I don't know" or "I should not act." In robotic harvesting, occlusion makes single-frame decisions fundamentally ambiguous: identical pixels can correspond either to a cuttable stem or to no stem at all. Existing VLAs are forced to commit, leading to high-confidence errors with irreversible consequences. We argue that the failure mode of a VLA is determined not by backbone scale but by its output interface. We propose Rejectable and Calibrated Decision Heads (RCDH), a typed, rejectable, and calibrated output interface that can be attached to a frozen VLA backbone without retraining or new features. RCDH introduces (i) a decision schema with explicit rejection and ordered, conditional decomposition, and (ii) a calibration procedure for risk-aware abstention. We evaluate RCDH on a robotic harvesting platform with controllable leaf occlusion, comparing generative, enumerated, calibrated, and rejectable interfaces. We show that replacing only the output head restores out-of-distribution usability under occlusion while preserving in-distribution performance. We further test whether the ordering of the rejection space is critical. Our results suggest that the right to refuse, rather than a larger model, is the missing interface for reliable manipulation under uncertainty.