Papers for

household robotics developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

HarnessPAI improves robot behavior by evolving executable programs

HarnessPAI: An Evolving Harness for Physical AI

Abstract: Physical AI aims to build embodied agents that perceive the world, understand and reason about it, and decide how to act. Yet the field has focused primarily on the last component: the action model that maps observations to low-level controls. The prevailing training recipe can erode the perceptual and reasoning capabilities needed for robust behavior, leaving even strong action models vulnerable to scene perturbations and long-horizon tasks. We introduce HarnessPAI, a model- and embodiment-agnostic Harness framework for Physical AI that treats code as the executable and evolvable interface that organizes the underlying action primitive. The framework separates two timescales: within a rollout, it executes open-loop at the program level, with a fixed program guiding and checking execution; across rollouts, it evolves closed-loop, using execution feedback to revise the program and distill failures into reusable skills. Across desktop robot arms, household robots, a robot vacuum, and a legged walking agent, HarnessPAI improves on both pure action models and code-as-policy baselines without retraining the underlying model: a 61.6-point gain over $π_{0.5}$ on LIBERO-PRO and a 27.2-point gain over WorldDreamer on RoboCasa atomic tasks. Once a program is selected, rollout execution requires no online high-level LLM deliberation. Beyond execution, the converged program is also a cheap and reliable expert-data collector, and fine-tuning $π_{0.5}$ on collected expert data lifts success rate on LIBERO-PRO by 38.8 points. Our results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system. Website: https://darwin-agent.github.io/HarnessPAI

Thu 24 SeptRoboticsArtificial Intelligence
The gist
Physical AI tries to build robots that can sense, think, and act in the world, but most work focuses only on making the robot move correctly. The authors introduce HarnessPAI, a framework that uses computer programs to guide and improve robot actions over time by learning from success and failure. This approach works across different robots and tasks without retraining the underlying action model, making robots more reliable and adaptable. The programs also help collect expert data that can further improve robot performance.
Open → 2609.29166v1

Vague2Detect improves detection of ambiguous household object prompts

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

Abstract: Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World frequently misalign vague language with the intended objects. Building on our prior work Commonsense-Guided Open-World Object Detection Using LLMs and Visual-Semantic Matching, we address YOLO-World's limitations in grounding task-driven queries. We propose Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base (KB), and YOLO-World verifies their presence in the image. For prompts outside the KB, a large language model (GPT-3.5-turbo) generates candidate descriptions, dynamically expanding the KB to cover novel concepts. On a benchmark of household scenes using custom images and an Open Images V7 subset, YOLO-World alone achieves only 32% Vague Prompt Success Rate (VPSR), the ability to map ambiguous queries to correct detections. In contrast, Vague2Detect improves performance to 61% VPSR with high precision, and up to 85% when augmented with GPT fallback.

Wed 9 SeptComputer Vision and Pattern RecognitionComputation and LanguageMachine Learning
The gist
Detecting objects in images can be tricky when people use vague or unclear descriptions. The authors show that popular models like YOLO struggle to recognize such vague prompts correctly. They created Vague2Detect, a system that uses a mix of language understanding and a knowledge base to better match ambiguous prompts with objects seen in images. This approach greatly improves the accuracy of detecting objects from unclear queries and can even handle new descriptions using a large language model.
Open → 2609.09949v1