Robot manipulation improves with task-focused local visual features

TLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot Manipulation

Robotics

Summary

Robots that follow language commands often struggle because the visual clues they need for each task get mixed up inside complex processing layers. The authors improved a robot control system by adding a clear way for the robot to focus on local visual details that match the task. This helps the robot understand and adapt to changes like different camera angles or noisy sensors, making it much better at completing a variety of tasks. Their method also lets people see exactly what parts of the scene the robot is paying attention to.

What this means in practice

  • For robotics engineers: Create multitask robot control systems that handle visual scene changes with clearer task-specific visual signals.
  • For industrial automation teams: Deploy robots in dynamic environments where visual conditions vary, improving success rates under camera and sensor changes.

Authors

Xianbo Cai, Hideyuki Ichiwara, Zihang Wang, Yijun Lu, Tetsuya Ogata

Abstract

Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing explicit, task-aligned local visual channel. We present TLC-DiT, a plug-in extension of the Multitask Diffusion Transformer (DiT) policy that adds explicit task-guided local visual feature maps without changing the diffusion objective or the action-generation process. For each camera view, frozen DINOv2 patch features are modulated by the CLIP task embedding through FiLM and refined by a lightweight CoordConv CNN adapter into smooth spatial maps, which are concatenated with the original global image, language, joint-state, and timestep conditions. On LIBERO, TLC-DiT reaches a 93.5% average success rate, compared with 86.5% for Multitask DiT and 79.25% for SmolVLA. On LIBERO-plus, the total success rate improves from 54.07% to 57.24%, with larger gains under camera, background, and sensor-noise changes. In real-world bimanual tasks, TLC-DiT raises Teabag Putting completion from 44% to 89% while maintaining comparable Match Box Opening performance. Feature-map visualizations confirm that the model attends to task-relevant regions across views and perturbations, providing a direct way to inspect the visual evidence.