Two-stage power sampling improves large vision-language model reasoning
ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts
Computer Vision and Pattern Recognition
Summary
Large vision-language models (LVLMs) try to answer questions about images and text but exploring all possible reasoning paths can be difficult. The authors show that a new two-stage approach called ReSight-SMC lets these models better explore different reasoning steps and image parts without extra training. This method groups possible answers and the way they were reached to reduce conflicts and improve final answers. Tests show ReSight-SMC works better than previous sampling approaches and is competitive with models trained by more complex methods.
What this means in practice
- •For machine learning engineers: Improve multimodal reasoning in vision-language applications using a training-free sampling approach that enhances answer diversity and accuracy.
- •For computer vision developers: Enhance image-grounded response generation by directing model attention to relevant visual regions during inference, improving answer quality without retraining.
- •For chatbot developers: Generate more accurate multimodal answers in conversational AI by sampling diverse reasoning paths supporting the same answer from vision-language models.$Commercial implications: Enables creation of improved vision-language chatbots with better reasoning skills, offering a competitive product in AI-assisted conversation tools.
Authors
Yaowen Zhang, Xiangyu Qiu, Junyi Hu, Zhi Lu, Wenwen Tian, Aoqin Wang, Junhai Luo, Zhenming Peng
Abstract
Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We transfer Power-SMC to LVLM decoding by defining a sequence-power target conditioned on both the image and the prompt. This direct transfer provides a strong training-free baseline, but leaves two aspects of finite-particle multimodal inference unaddressed. At the particle level, global resampling can collapse genealogies, while particle-based power sampling does not diversify trajectories through distinct visual cues in multimodal decoding, limiting exploration under a finite particle budget. At the answer level, sequence-level sharpening makes distinct reasoning trajectories compete even when they support the same answer. We introduce ReSight-SMC, a verifier-free two-stage power sampler for LVLM inference. Its first stage uses ancestry-isolated SMC islands to preserve independent trajectory families and routes a bounded set of prefix-conditioned visual scouts to prefix-relevant image regions while discouraging redundant overlap. Each scout temporarily increases attention to the image tokens and emphasizes its routed region. Exact importance correction preserves the base LVLM sequence-power target. The second stage aggregates terminal importance mass by canonical answer, powers the answer marginal, and samples an answer together with a supporting trajectory. Across four LVLM backbones and five benchmarks, ReSight-SMC achieves stronger aggregate performance than Power-SMC over both the reasoning and perception benchmark groups. Without post-training, it remains competitive in aggregate with backbone-matched models trained using reinforcement learning.