Papers for

interactive game developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Sub-goal guided reflection improves long task success for language agents

SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning

Abstract: Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLACT, which reflects only on the end-goal at each step without explicitly considering intermediate sub goals. To address this problem, we propose SGG-ReflAct (Sub-Goal Guided Re flAct), a reasoning backbone that integrates sub-goals generated through a single path LLM planner into the reflection process. We further extend this framework to BeamSGG-ReflAct, which replaces the single-path planner with a beam search based LLM planner for structured plan exploration. We run experiments on ALF World, ScienceWorld, and Jericho with multiple LLM models. SGG-ReflAct out performs REFLACT in nearly all settings, achieving best success rate gains of 14.9 percentage points on ALFWorld and 8.0 percentage points on ScienceWorld with Llama-3.1-8B-Instruct. Our experimental analysis shows that SGG-ReflAct re duces hallucinated actions and achieves its largest gains on procedurally ordered tasks. Furthermore, experimental results with BeamSGG-ReflAct show that the backbone's effectiveness depends on plan quality: explicitly specifying the re quired operations recovers gains that plan searching alone cannot achieve. These results demonstrate that SGG-ReflAct offers a practical and highly effective rea soning backbone, enabling LLM agents to achieve reliable performance in com plex, long-horizon tasks through easy integration.

Mon 28 SeptArtificial Intelligence
The gist
Large language models struggle with long, complex tasks because they can lose track of important steps along the way, leading to mistakes. The authors propose a new approach that breaks down big goals into smaller sub-goals and guides the model to reflect on these intermediate steps. This method helps the model stay on track and complete complicated tasks more reliably, showing big improvements in several test environments. The technique also explores multiple plans to find better strategies, making language agents better at handling long sequences of actions.
Open → 2609.34548v1

Mamba 3 model updated to track order sensitive states better

Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3

Abstract: State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations matters. The rank-one update couples state coordinates along an input-dependent direction, enabling non-diagonal state transitions within a single Mamba-3 block. The extension preserves Mamba-3's exponential-trapezoidal discretization, rotary embeddings (RoPE), and readout. For training, we adapt chunkwise computation to parallelize the proposed recurrence within each chunk. Experiments cover group word problems with discrete inputs and a shell game with continuous observations, in which a policy is trained by behavioral cloning. Among the models selected for their strong performance under fixed timing, the proposed model maintains higher tracking success on longer swap sequences in the shell game with continuous observations and timing jitter. These experiments show that the proposed method achieves high accuracy on the evaluated non-commutative tracking tasks, improving on standard Mamba-3. The extension thus offers a Mamba-3-based approach to non-commutative state tracking.

Wed 23 SeptMachine Learning
The gist
Tracking the order of events is important in many tasks because the sequence affects the outcome. The authors improved a model called Mamba-3 by adding a new way to update its internal memory that depends on the input and links different parts of the memory together. This allows the model to better remember sequences where order matters. They tested it on challenges involving tracking swaps and games, showing it works better than before when things happen in complex orders. This makes it more reliable for tasks where the order of observations can't be mixed up.
Open → 2609.28273v1

Motion based captchas reveal human perception edge over gui agents

Invisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents

Abstract: Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened by advances in multimodal large language models (MLLMs) and Graphical User Interface (GUI) agents, which exhibit strong visual perception, reasoning, and browser interaction capabilities. We propose Motion Vision CAPTCHA (MVCAP), a hierarchical motion-based CAPTCHA framework in which target semantics are instantiated as motion-defined foreground structures and become recoverable only through temporal segregation from a dynamically evolving background. Built on this shared principle, MVCAP is instantiated in three perceptually progressive levels: coherent motion, structural motion, and biological motion. To evaluate this framework, we introduce MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances, together with a matched foreground-only control benchmark, MVCAP-Bench-FG. We evaluate humans, Browser Use agents, native computer use agents, and a supplementary offline VQA setting derived from the same instances. Results reveal a substantial human--agent gap: on the full MVCAP-Bench, human accuracy reaches 99.6%, whereas the best GUI agent achieves only 16.8%, close to the six-way chance level. The foreground-only control further shows that the key difficulty comes from dynamic background camouflage rather than answer format or browser interaction alone. These findings identify a measurable human--agent perception gap and position MVCAP-Bench as a benchmark for studying motion-defined perception in current agents.

Wed 23 SeptComputer Vision and Pattern Recognition
The gist
Many captchas are solved by looking at static pictures, but smarter computer programs can now recognize and click through these easily. The authors created a new kind of captcha that relies on moving images, making it easy for humans but very hard for current computer agents to solve. This happens because the important parts of the image only appear clearly when watching how things move over time, rather than from a single snapshot. Their tests showed humans scored nearly perfect, while computer agents barely did better than guessing.
Open → 2609.27461v1

Gesturefar enables real-time natural gesture generation from streaming speech

GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

Abstract: Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality--latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: https://andypinxinliu.github.io/GestureFAR

Fri 18 SeptComputer Vision and Pattern RecognitionGraphicsHuman-Computer Interaction
The gist
Making digital characters move their hands naturally while talking is tricky, especially when the speech is still happening. Previous methods simplified complex hand motions into limited building blocks, which made the gestures less lifelike and varied. The authors created GestureFAR, a new way that predicts smooth, continuous hand movements directly from what is being said, without waiting for the speaker to finish. This approach lets virtual agents produce more natural gestures instantly as speech flows.
Open → 2609.21576v1

Self emergence architecture creates distinct stable agent personalities

Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling

Abstract: Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, maintaining neither behavioral inertia nor endogenous self-evolution. We propose the Self-Emergence Agent Architecture (SEAA), which integrates three components: (i) a Hidden Markov Model (HMM) that encodes long-term behavioral and cognitive inertia as an editable state-transition matrix; (ii) a Reflexion-style verbal metacognition loop whose output updates the HMM parameters themselves, rather than merely being stored as text; and (iii) a multi-agent social environment in which initially identical agents continuously compare their behavior with others'. The three components form a closed loop: social action $\to$ feedback $\to$ self-reflection $\to$ inertia update $\to$ differentiated action. We state three falsifiable hypotheses and provide a reproducible experimental protocol with operational metrics. A language-model-free prototype shows the loop spontaneously breaks symmetry: initially identical agents consolidate distinct, stable personalities whereas matched controls do not. Experiments with a hosted LLM surface these differences as distinct first-person self-narratives, and a five-agent deliberation spontaneously develops social structure---a consensus hub and a unanimously rejected outlier---absent in the control. Following an epistemologically agnostic stance inspired by Zhuangzi, SEAA studies only observable behavioral emergence and makes no claim about subjective qualia. This work contributes a unified framework, a concrete architecture with pseudocode, mechanistic evidence, and a microscope-style sandbox for studying artificial-self emergence.

Tue 15 SeptArtificial Intelligence
The gist
Language model agents can generate language well but often struggle with keeping a consistent personality and recognizing themselves versus others. The authors created a new system where agents update their behavior based on past actions, think about their own thoughts, and compare themselves with other agents. This loop helps initially identical agents develop stable, different personalities and social roles. Their experiments show this method forms clear social structures and distinct self-narratives that previous designs do not produce.
Open → 2609.17331v1

Simulation shows how personas and models shape story character choices

ANIMASK: What the Model Contributes to Role Play in Simulated Story Worlds

Abstract: When a language model plays a character, the observed behavior reflects both the assigned persona and the default dispositions of the actor model itself. Existing evaluations test persona fidelity or model defaults in isolation, but neither says, at a specific choice with consequences, what the persona changed and what the model's default kept. We introduce ANIMASK, a simulation framework that freezes books and scripts into story worlds whose characters act on their own motivations and replays each story from its freeze point. We hold out the author's continuation as a human reference, verify through in-story interviews that each persona remains present, and at every decision point compare the character's action with what the model produces when the persona is removed. Across 40 stories, 6 actor models, and 3,846 decision points, the replays converge away from their canons in one shared direction, toward flatter, cooler stories that leave their tensions open. The personas stay present and obeyed throughout. On three choices in four the model's default already falls inside what the persona accepts, and where the two diverge the model is the cautious one, holding where the persona would press. The persona guarantees who the character is, and the model sets how far the character will go.

Tue 15 SeptArtificial Intelligence
The gist
When a language model plays a story character, the character's actions come from both the personality given to them and the language model's own tendencies. The authors created a system called ANIMASK to freeze stories and replay them, seeing where the model's default choices differ from the character's assigned persona. They found that while the personas remain clear, the models tend to make safer, less tense choices than the original stories. This means the persona defines who the character is, but the model decides how adventurous or cautious that character acts.
Open → 2609.16667v1