Vague2Detect improves detection of ambiguous household object prompts

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

Computer Vision and Pattern RecognitionComputation and LanguageMachine Learning

Summary

Detecting objects in images can be tricky when people use vague or unclear descriptions. The authors show that popular models like YOLO struggle to recognize such vague prompts correctly. They created Vague2Detect, a system that uses a mix of language understanding and a knowledge base to better match ambiguous prompts with objects seen in images. This approach greatly improves the accuracy of detecting objects from unclear queries and can even handle new descriptions using a large language model.

What this means in practice

  • For household robotics developers: Enable robots to interpret vague user commands about household items more accurately by linking ambiguous descriptions to known objects in their environment.
  • For smart home device makers: Improve smart devices’ capability to recognize user references to household objects even when users provide unclear or functional descriptions.

Authors

Ibrohimjon Muminov, Jihie Kim

Abstract

Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World frequently misalign vague language with the intended objects. Building on our prior work Commonsense-Guided Open-World Object Detection Using LLMs and Visual-Semantic Matching, we address YOLO-World's limitations in grounding task-driven queries. We propose Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base (KB), and YOLO-World verifies their presence in the image. For prompts outside the KB, a large language model (GPT-3.5-turbo) generates candidate descriptions, dynamically expanding the KB to cover novel concepts. On a benchmark of household scenes using custom images and an Open Images V7 subset, YOLO-World alone achieves only 32% Vague Prompt Success Rate (VPSR), the ability to map ambiguous queries to correct detections. In contrast, Vague2Detect improves performance to 61% VPSR with high precision, and up to 85% when augmented with GPT fallback.