Papers for

virtual reality designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Intentional agent method improves humanlike facial reactions in ai systems

Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation

Abstract: Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG methods rely on end-to-end architectures that directly map speaker behaviours to listener expressions without an explicit intermediate internal state. We reformulate FRG as generation mediated by a structured internal-state process and propose the \textbf{Intentional Agent}, which shifts FRG from direct stimulus-response mapping to stimulus-grounded generation through explicit intermediate states. To represent temporal internal-state evolution, we propose an internal dynamics model that integrates emotional drives with an iterative Inner Thought Flow (ITF) within a structured intermediate state used for subsequent generation. This state can continue to update during conversational silences. Furthermore, to bridge abstract internal states with physiological actions, we formulate FRG as a downstream affective mapping from this latent thought flow to facial expressions. Experiments on the REACT 2025 dataset show an FRDist of 72.39 and an FRDiv of 0.5057; perceptual plausibility is evaluated separately through blinded human ratings. A blinded human evaluation of 96 reactions found no significant difference in mean score between Full and ground truth ($5.527$ vs.\ $5.195$, $p_{\mathrm{Holm}}=.076$), while Full significantly outperformed Event-Triggered and Heuristic-Only (both $p_{\mathrm{Holm}}<.001$). The Reaction Quality Scorer (RQS) correlated strongly with human judgements (Pearson $r=.855$; Spearman $ρ=.821$, both $p<.05$), supporting its use as an automatic metric. These results underscore the immense potential of endogenous dynamics in building highly autonomous, human-like agents.

Mon 28 SeptArtificial Intelligence
The gist
Making computers show natural facial reactions is important for better conversations with humans. The authors propose a new method that makes AI agents think internally before showing expressions, rather than just copying what they see. Their approach uses a model that updates an inner emotional state even during pauses, leading to more realistic and timed facial reactions. Tests with human judges showed their system’s expressions were as believable as real ones, better than some older methods.
Open → 2609.34419v1

FloodDiffusion 2 speeds up and controls streaming motion generation

FloodDiffusion 2: Efficient and Path Controllable Streaming Motion Generation

Abstract: We present FloodDiffusion 2 (FD2), an efficient and controllable framework that builds upon FloodDiffusion (FD1), a state-of-the-art streaming motion generation model. While FD1 produces plausible motion, it suffers from low efficiency and limited controllability, as its attention design requires repeated computation over the entire history, and it lacks precise trajectory control for real-world applications. To address these limitations and improve generation quality, FD2 introduces three advances. First, Partial Attention makes finalized history representations independent of the active window, enabling KV-cached inference and shared-history packing for efficient training. Second, we establish a necessary-and-sufficient Bregman criterion for regression losses to preserve diffusion's conditional-mean velocity field. This criterion guides an FK-induced quadratic loss that incorporates motion geometry without online FK evaluation. Third, FD2 introduces precise path conditioning to control the character's root trajectory while preserving natural body motion. Experiments show that FD2 reduces training computation by 4.6$\times$ and accelerates denoising by 11.29$\times$, reaching 2.303 ms per update on long sequences. Alongside these efficiency gains, FD2 improves motion quality over FD1 and achieves state-of-the-art FID scores among streaming methods, with 0.048 on SEED and 0.053 on HumanML3D.

Sun 27 SeptComputer Vision and Pattern Recognition
The gist
Generating smooth and natural human-like motion on computers can take a long time and be hard to control precisely. The authors improved an existing system, FloodDiffusion, by redesigning how it remembers past movements and adding a way to guide the character’s path exactly. Their method makes generating motion much faster—over ten times quicker during key steps—and produces better, more realistic movements. This helps create smoother animations for long sequences, useful in games or virtual environments.
Open → 2609.33167v1

Human motion created by combining spatial audio and text intent

MoSAT: Human Motion Generation from Spatial Audio and Textual Description

Abstract: Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the desired action and how it should be performed. In this paper, we study the novel task of human motion synthesis jointly conditioned on spatial audio and natural language, a problem that has been largely overlooked in previous research. To support this task, We introduce STAM, a dataset of motion sequences paired with spatial audio and detailed textual annotations whose rich vocabulary affords precise and nuanced specification of human motions. We further introduce MoSAT, a latent flow-matching framework for full-body motion generation jointly conditioned on natural-language intent and directional spatial-audio cues through hierarchical cross-attention before generating motion. Such a hierarchical design enhances temporally coherent and semantically aligned motion sequences. We also develop tri-modal evaluators for comprehensive evaluation on this novel task. Extensive experiments show that MoSAT achieves the SOTA performance by leveraging spatial audio's intrinsic motion-shaping properties alongside textual semantics, enabling precise and diverse motion in various scenarios.

Sun 20 SeptGraphicsComputer Vision and Pattern RecognitionRobotics
The gist
People's movements are influenced by what they hear around them and what they want to do. This paper introduces a way to create human-like full-body motions using both sounds coming from specific directions and written descriptions of actions. The researchers also made a new dataset that matches motion data with spatial sounds and detailed text, helping their system learn better. Their method, called MoSAT, improves how natural and accurate the generated movements are by linking sounds, language, and motion smoothly. They also created special tools to check that the motions match the audio and text properly.
Open → 2609.23797v1

Physics grounded system synthesizes dynamic multi part object scenes

PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians

Abstract: Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with the Material Point Method (MPM) to generate physically driven motion, but extending this paradigm to heterogeneous multi-part objects and interacting multi-object scenes remains challenging. Object-level physical assignment collapses distinct parts into a single material state, while one-shot predictions from large language models, vision-language models, or agents neither reliably bind different materials to identified parts nor verify that the resulting MPM configuration is executable. Score Distillation Sampling (SDS)-based parameter optimization, meanwhile, requires repeated per-scene score evaluations and gradient backpropagation, incurring lengthy optimization and potentially yielding suboptimal or unstable solutions. We therefore present PhysMAS, a physics-grounded multi-agent framework. From a motion prompt and four scene views, an Object-Part Scene Agent establishes persistent identities and calls a Material Reasoning Agent for part-wise profiles. It invokes solver-aware skills to bind these identities and profiles to per-particle MPM fields and execute all objects in a shared domain; the framework then screens candidate forward-simulation results. This supports heterogeneous multi-part and interacting multi-object scenes without per-scene diffusion-score backpropagation. Extensive experiments demonstrate that, compared with recent physics-based 4D Gaussian baselines that rely on SDS, PhysMAS achieves better semantic alignment and perceived physical plausibility while requiring less runtime.

Mon 7 SeptArtificial Intelligence
The gist
Creating believable animations of moving objects made of many parts is hard, especially when those parts interact physically. The authors present PhysMAS, a system that uses multiple agents to understand and assign physical materials to different parts in a scene, ensuring realistic motion according to physics. Unlike previous approaches, PhysMAS avoids slow optimizations and handles scenes with multiple objects interacting naturally. This leads to more physically plausible and semantically accurate dynamic scenes faster.
Open → 2609.07174v1