Working memory distillation improves reasoning in image segmentation models

Revisit to Segment: Working Memory Distillation for Reasoning Segmentation

Computer Vision and Pattern Recognition

Summary

Image segmentation means figuring out where things are in pictures. Some smart computer models can think about pictures and explain their reasoning step-by-step. This paper found that when these models remember their earlier thoughts, they get better at understanding images. The researchers created a way to teach simpler models by using these memories, so even without remembering, the simpler models perform better. Their new technique, called SWiM, was tested and showed improved accuracy in segmenting images.

What this means in practice

  • For computer vision developers: Use working-memory distillation to improve image segmentation accuracy in models without extra memory at runtime.
  • For autonomous vehicle engineers: Enhance object recognition and localization in driving scenes by deploying models trained with working-memory guidance for better segmentation.

Authors

Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu, Xiaolong Jiang, Yao Hu, Jiayin Cai

Abstract

Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our exploration reveals that MLLMs benefit from using this self-generated working memory as context, leading to enhanced reasoning segmentation. Motivated by this finding, we seek to strengthen the backbone model's reasoning segmentation capabilities by distilling the guidance gained from revisiting prior attempts, enabling it to benefit with or without working memory at inference time. To this end, we propose Reasoning Segmenter with Working Memory (SWiM), a working-memory distillation framework for reasoning segmentation. Specifically, SWiM selects rollouts based on segmentation quality to construct working memory and uses the memory-conditioned model as a teacher. The teacher provides token-level distributional supervision along student-generated trajectories, while the student receives only the original image and query. Joint optimization of on-policy self-distillation and outcome-based reinforcement learning combines working-memory guidance with direct feedback on segmentation quality. Extensive experiments on reasoning segmentation benchmarks demonstrate that SWiM achieves state-of-the-art performance, validating the effectiveness of working-memory distillation.