Papers for

smart home device makers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

MmWave radar enables full 3D body mesh without cameras

Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision

Abstract: Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-modal teacher-student framework that lifts commercial radar to full-body, per-frame, metric 3D mesh reconstruction with per-joint uncertainty. Three innovations: (1) a mesh-foundation-model teacher - SAM 3D Body produces whole-body MHR ground truth (70 joints, 18,439 mesh vertices) from a single RGB frame with zero training, slashing annotation cost by orders of magnitude; (2) StudentPoseFormer - set encoding with masked attention pooling, a temporal Transformer, and a CVAE multi-hypothesis head that outputs both the pose mean and per-joint variance, honestly reporting where the radar cannot see; and (3) a multi-stage ground-truth quality pipeline (confidence gating, depth validation, temporal smoothing, bone-length consistency, bad-frame rejection) plus systematic information-lever ablations. On the public MM-Fi benchmark (same TI IWR6843 sensor, cross-subject), our full configuration reaches 7.45 cm 12-joint MPJPE, with ablations proving the causal value of point accumulation (k = 3, -0.34 cm), Doppler (-0.85 cm; -2 cm at the wrist on fast actions), and velocity loss (-0.27 cm). On our own synchronized radar + RGB-D corpus with block-level held-out splits, the pipeline achieves 21.47 cm end-to-end (per-joint hierarchy from 4.8 cm at the hip to 34.7 cm at the wrist - matching physical information limits), could be improved to 15 cm with ~30k diverse samples, and a scaling law shows sample diversity, not volume, is the binding constraint. Deployment inference is radar-only - no camera, no image.

Mon 28 SeptArtificial IntelligenceComputer Vision and Pattern Recognition
The gist
Measuring people’s body shapes and movements often uses cameras that can invade privacy. This paper shows how to use very sparse radar data instead, turning it into detailed 3D body models without any cameras involved during operation. The authors use a smart teacher-student system where a model trained on camera images teaches another model to understand radar signals. Their new approach estimates not only body poses but also how uncertain each measurement is, improving reliability while keeping people's privacy intact.
Open → 2609.34768v1

Audio language models struggle to ignore irrelevant speakers near voice assistants

Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents

Abstract: Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch scenarios. Each item uses a shared action space comprising silence, a tool call, and a natural-language answer. Speaker-switch pairs hold the specified words fixed while source, distance rendering, and a temporal boundary define a controlled wearer-to-bystander shift. Six raw Audio LLMs and three training-free adaptations often identify the target tool yet rarely withhold action under this shift; the highest raw switch mute rate is 14%. We then use VoxGate as a post-training case study. Supervised training mutes 91.3% of switched commands while choosing the correct tool for all nearby wearer commands and text-only controls. An exploratory GRPO stage has similar switch performance; side-talk accuracy rises from 68.4% to 70.9%, and self-talk muting from 52.0% to 60.0%. Factorized controls identify an independent source-change effect, while sensitivity to the far-field manipulation varies across acoustic renderings. The benchmark therefore measures multi-cue acoustic-context gating rather than isolated speaker identity.

Sat 26 SeptSoundArtificial IntelligenceComputation and Language
The gist
Voice assistants need to decide when to respond to spoken commands, especially when many people talk nearby. The authors created a test called VGBench to check if audio language models can tell when to act or stay silent based on who is speaking and the situation around them. They found that current models often respond even when they shouldn't, like reacting to a bystander instead of the user. The authors improved this by training a system called VoxGate, which helps the assistant correctly ignore commands from others while still responding to the user.
Open → 2609.32536v1

Omni-language models enable zero-shot audio visual navigation

RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation

Abstract: We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new \emph{Global Navigation Instruction} setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Code: https://github.com/rikeilong/OmniAV\_Nav.

Sat 26 SeptArtificial Intelligence
The gist
Navigating by combining what you hear and see is hard for robots, especially without training on specific tasks. The authors developed RAO-Nav, a method that uses large language models capable of understanding both sounds and visuals to help robots reason and move around new places without prior training. They also created a way to improve the robot's navigation decisions by encouraging it to focus on the most relevant information. Their approach works better than specialized models on existing tests without any extra training. They also suggest a new challenge to test how well these models can handle broader navigation instructions.
Open → 2609.32224v1

Speaker distance estimates improve with few real labeled examples

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Abstract: Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.

Thu 24 SeptSound
The gist
Measuring how far someone is from a microphone is tricky because models trained on simulated sounds don't work well with real recordings. The authors found that guessing the average distance is often better than using these models directly. However, by using just a small number of real recordings with known distances, they can adjust the model outputs to get better results without retraining. This process focuses on keeping the order of distances correct rather than perfect accuracy. Their work helps decide how many real examples are needed to improve these distance estimates effectively.
Open → 2609.29203v1

Robots learn to take initiative and help without being told

Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots

Abstract: Effective robot assistance beyond narrow roles and repetitive tasks requires robots to be proactive - to decide what needs to be done rather than waiting to be told. While proactivity is increasingly explored, it lacks a unified formulation, and work in the domain is typically evaluated offline against static human models that cannot capture the effect of a robot's actions on the environment and the user's own behavior. We introduce a unified formalism for proactive robot assistance, organize it into three levels, and provide a framework to address the highest level of unprompted proactive assistance. We then show that offline evaluation overstates performance in this setting, and contribute a closed-loop evaluation with a human model that adapts to the robot. Finally, we present a method, GAP, that instantiates our framework, learning from passive observation to anticipate user goals and act. Under closed-loop evaluation, prior state-of-the-art methods collapse, in some cases adding more work than they save, while GAP remains robust and substantially outperforms them.

Thu 24 SeptRoboticsArtificial Intelligence
The gist
Robots often wait to be told what to do, which limits how helpful they can be. This paper introduces a way to make robots more proactive, letting them figure out what needs doing on their own. The authors show that testing these proactive robots by just simulating people can be misleading, so they created a better test where the robot’s actions influence how the human behaves. They also developed a new method called GAP that learns from watching people and can predict what they want, helping more effectively than past approaches.
Open → 2609.28910v1

WiFi sensing method separates motion speed from location effects

Untangling the Geometry and Speed for RF Sensing Spectrograms

Abstract: A fundamental challenge in RF sensing is that Doppler signatures observed by a link entangle the target's motion with the sensing geometry, resulting in limited applicability to unconstrained real-world settings. In this paper, we establish a new foundation for physically interpretable RF sensing that disentangles reflector speed from geometry, jointly recovering the speed, geometry factor, relative amplitude, and width of each dominant Doppler ridge. More specifically, we first develop a compact parametric representation of WiFi spectrograms and establish its low-dimensional structure through a systematic computer-vision analysis of a large and diverse human-activity dataset, thereby providing a tractable foundation for learning. Building on this representation, we then design a physics-informed autoencoder whose structured bottleneck and differentiable RF forward model enforce physically meaningful estimates of reflector speed and geometry. We further introduce a synthetic-to-real training framework, eliminating the need for real WiFi training data. We extensively validate the proposed framework under both known and time-varying geometries, using both independently generated synthetic test sets and 31 real WiFi experiments. The results demonstrate the superior performance in speed and geometry extraction, robustly recovering the underlying geometry, speeds, Doppler-ridge amplitudes, and ridge widths across all settings, while substantially outperforming the strongest baselines.

Tue 22 SeptMachine Learning
The gist
WiFi signals can be used to sense movement, but the signals often mix up how fast something moves with where it is. The authors created a new way to separate these two effects so that the sensing is clearer and more accurate. They developed a model that looks at WiFi signal patterns and can figure out both the speed and position of a moving object. Their system was tested with many real and simulated examples and outperformed older methods.
Open → 2609.26960v1

Robot learns to help find objects using language and actions

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

Abstract: Robot assistants for older adults and people with disabilities need to perform collaborative tasks with users effectively. The core component of these systems is an interaction manager whose job is to observe and assess the task and infer the state of the human and their intent for the robot to choose the best course of action. Due to the sparseness of the data in this domain, the policy for such multimodal systems is often crafted by hand; as the complexity of interactions grows, this process is not scalable. This paper proposes a reinforcement learning (RL) approach to automatically generate the multimodal policy of the robot. Our system focuses on a realistic scenario where a robot assists a user in locating objects within a home environment, managing multimodal signals, including language and physical actions, to select the best action. In contrast to traditional dialog systems, our agent is trained with a simulator that uses human data and can deal with multiple modalities. We use a simple high-level reward function that needs no fine-tuning and enforce some preconditions to speed up the training process. A human study evaluating the system in a real-world setting demonstrates promising results, indicating high usability and effective task completion. This RL-based approach offers a scalable and interpretable alternative for designing interaction managers in multimodal human-robot collaborations.

Mon 21 SeptRobotics
The gist
Helping robots assist people at home, especially older adults or those with disabilities, is hard because robots must understand what humans want through speech and body movements. The authors show a way for robots to learn how to plan their actions automatically by practicing in a virtual setup that mimics real human behavior. Their method lets robots decide the best thing to do next without people programming every step. When tested with real users, their approach helped robots work well and made tasks easier to complete.
Open → 2609.25274v1

GPT-6-Astra shows strengths and limits in zero-shot robot navigation

GPT-6-Astra in a Navigation Workflow: Behavioral Analysis in Zero-Shot Vision-and-Language Navigation in Continuous Environments

Abstract: We study GPT-6-Astra in a zero-shot Vision-and-Language Navigation in Continuous Environments (VLN-CE) system, where it interprets instructions, assesses its surroundings, and proposes actions. The system uses a common observation--decision--execution workflow with direct model API calls, without a packaged agent harness or navigation-specific fine-tuning. In this workflow, each request receives selected observations, execution feedback, and retained progress records. Evaluation covers the complete system, including context management and action control. We evaluate the system on 50 of the 100 R2R-CE val-unseen episodes used by Open-Nav. It achieves a success rate of 52.0\%, an SPL of 48.9\%, and an nDTW of 70.8\%. Our analysis highlights three findings. First, recorded responses link landmarks and earlier actions to instructions using observations and supplied history. Second, reviews include requests for additional views and revisions of uncertain judgments. Third, the results suggest a gap between task understanding and autonomous completion: an unfinished crossing is recognized while rotation continues. At termination, 36.0\% of episodes succeed with a workflow-accepted STOP, while another 16.0\% meet the distance criterion at the step limit. These results highlight a central challenge: translating correct local judgments into sustained progress and appropriate stopping.

Thu 17 SeptRobotics
The gist
This study looks at how GPT-6-Astra can understand spoken or written directions and move through a space it has never seen before, without prior training for navigation. The authors found that the system can describe landmarks and connect past moves to instructions, sometimes asking for better views before deciding what to do. However, the system struggles to stop at the right time or keep making progress after recognizing problems, showing a gap between understanding the task and completing it. This means GPT-6-Astra can make smart local decisions but has trouble finishing navigation tasks fully on its own.
Open → 2609.20116v1

Vague2Detect improves detection of ambiguous household object prompts

Vague2Detect: Handling Ambiguous Prompts in Knowledge-Based Open-World Detection

Abstract: Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World frequently misalign vague language with the intended objects. Building on our prior work Commonsense-Guided Open-World Object Detection Using LLMs and Visual-Semantic Matching, we address YOLO-World's limitations in grounding task-driven queries. We propose Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base (KB), and YOLO-World verifies their presence in the image. For prompts outside the KB, a large language model (GPT-3.5-turbo) generates candidate descriptions, dynamically expanding the KB to cover novel concepts. On a benchmark of household scenes using custom images and an Open Images V7 subset, YOLO-World alone achieves only 32% Vague Prompt Success Rate (VPSR), the ability to map ambiguous queries to correct detections. In contrast, Vague2Detect improves performance to 61% VPSR with high precision, and up to 85% when augmented with GPT fallback.

Wed 9 SeptComputer Vision and Pattern RecognitionComputation and LanguageMachine Learning
The gist
Detecting objects in images can be tricky when people use vague or unclear descriptions. The authors show that popular models like YOLO struggle to recognize such vague prompts correctly. They created Vague2Detect, a system that uses a mix of language understanding and a knowledge base to better match ambiguous prompts with objects seen in images. This approach greatly improves the accuracy of detecting objects from unclear queries and can even handle new descriptions using a large language model.
Open → 2609.09949v1

Distributed microphones improve acoustic scene understanding with geometry

Geometry-Informed Distributed Acoustic Scene Understanding

Abstract: Acoustic scene understanding in multi-room environments is a difficult task. Most existing systems use a single centralized microphone array, and they often fail because walls and doors block sound signals. To address this challenge, we propose a geometry-informed distributed acoustic scene understanding framework. Our system leverages distributed microphones and uses an audio spectrogram transformer and a topology-aware graph neural network to fuse spatio-temporal acoustic features. Then, these features are decoded into discrete semantic triplets. Finally, a frozen large language model combines these symbolic observations with the environmental geometry. This allows the system to perform spatial understanding, infer plausible missing transitions, and generate a physically consistent narrative of the scene. Experiments on a custom multi-room simulator demonstrate that our framework outperforms centralized baselines and improves spatial consistency under simulated occlusion.

Mon 7 SeptSound
The gist
Understanding sounds across multiple rooms is hard because walls block sound signals. The authors propose a system that uses microphones placed in different rooms and combines their data using special neural networks and a language model aware of room layouts. This approach helps fill in missing information and creates a consistent story of what is happening in the space. Tests in simulated multi-room environments show this method works better than just using one microphone array.
Open → 2609.08026v1