Ego forge creates first person videos from third person footage

Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation

Computer Vision and Pattern Recognition

Summary

Seeing what a person sees just from videos taken by others is very hard because the viewpoint changes a lot and some things are never seen before. The authors developed Ego-Forge, a system that can guess what someone would see from their own eyes given third-person video, without needing extra instructions or special text captions. It learns from many videos and uses a clever new way to think about the scene without relying on complex geometry rules. This makes Ego-Forge faster, more general, and better at imagining new views than previous methods.

What this means in practice

  • For virtual reality developers: Generate realistic first-person views from third-person footage to enhance immersive VR experiences without complicated inputs.
  • For film post-production teams: Create egocentric video shots from third-person takes to expand creative options and viewpoints in editing workflows.

Authors

Mohammad Mahdi, Luc Van Gool, Danda Pani Paudel

Abstract

Exo-to-egocentric video generation aims to synthesize what a person sees from their own viewpoint given third-person footage and a target head trajectory. The task requires transferring appearance and semantics across large viewpoint changes while hallucinating content never observed by the exocentric camera. Existing approaches either impose additional input requirements, such as a ground-truth initial egocentric frame or multiple synchronized exocentric views, or remain limited to category-specific settings. EgoX is the first to address cross-activity and in-the-wild generalization, but requires a human-provided caption of the non-existent egocentric view at inference and introduces a computationally expensive geometry-guided attention bias that can propagate reconstruction errors and suppress textual and visual context. We therefore propose \textbf{Ego-Forge}, a caption-free and bias-free framework for exo-to-egocentric generation. It introduces \textit{Dynamic Captioning}, which derives conditioning tokens directly from the model's hidden states and adapts them to the diffusion timestep and network depth, replacing external text conditioning. By scaling training by an order of magnitude and using all available exocentric viewpoints, Ego-Forge learns cross-view correspondence implicitly and eliminates the need for geometry-guided attention, requiring only a lightweight depth prior. Ego-Forge achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearance inference. Our model and source code will be made publicly available.