Llms handle multiple tasks at once for smarter real time use

LLMs are General Asynchronous Agents

Machine LearningComputation and Language

Summary

Most language models now work by doing one step at a time: listen, think, then respond. But many real-world tools like voice helpers or robots get new information all the time and need to work on many things at once. The authors show how some large language models, like Qwen 3.x, can work in an asynchronous way, juggling multiple tasks simultaneously without needing extra special training. This helps machines understand things like videos or games while still managing other tasks at the same time.

What this means in practice

  • For voice assistant developers: Build voice assistants that process new conversation inputs while generating responses, improving real-time interaction quality.
  • For robotics engineers: Enable robots to handle multiple sensory inputs and control signals simultaneously without task-specific training.

Authors

George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov, Vladimir Kaurkin, Vadim Pastushenko

Abstract

Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.