Video dataset introduces detailed timing for better video and language models
Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Understanding videos often requires looking at detailed changes over time, but most datasets only give rough information. The authors created Kairos, a collection of long videos with precise notes about what is happening at each moment. These notes include actions, who is present, what attributes things have, and how contexts change as the video plays. This dataset helps computers learn and understand videos in a more detailed and continuous way over longer periods. It can be used to improve machine understanding, language instructions about videos, and even video creation.
video datasetvideo-language modelingtemporal alignmentfine-grained annotationsvisual dynamicslong-duration videosaction recognitioncontextual cuesrepresentation learningvideo generation
Authors
Ruibo Ming, Lei Sun, Deheng Zhang, He Zhang, Jialu Li, Jian Wang, Zhendong Li, Mengshun Hu, Danda Pani Paudel, Luc Van Gool, Jinjin Gu
Abstract
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.