CameraAnything: Refilming Videos with Arbitrary Camera Control
2026-07-27 • Computer Vision and Pattern Recognition
Computer Vision and Pattern Recognition
AI summaryⓘ
The authors present CameraAnything, a new method that lets users edit videos by changing both how the camera moves and its settings like zoom and resolution together. Unlike past methods that needed complex 3D models or only changed camera angles, their approach uses a special way to represent camera rays and positional data, making these edits more natural. They also created a synthetic dataset with multiple camera views to train their system well. This allows users to smoothly adjust camera viewpoint, zoom, and video size all in one go, which could help in making movie edits or adapting videos for different devices.
intrinsic camera parametersextrinsic camera parametersPlücker ray3D RoPEself-attentionvideo editingsynthetic training datacamera viewpoint controlfocal lengthresolution adaptation
Authors
Yixuan Li, Yanhong Zeng, Ka Leong Cheng, Jiayi Zhu, Hanlin Wang, Wen Wang, Yihao Meng, Hao Ouyang, Qiuyu Wang, Yue Yu, ZiDong Wang, Yiyuan Zhang, Yujun Shen, Dahua Lin
Abstract
We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters. Existing approaches either rely on expensive 3D reconstruction to achieve full camera functionality or restrict editing to extrinsic parameter manipulation. Moreover, the coupled influence of intrinsic and extrinsic parameters on video appearance makes disentangled modeling particularly challenging. To address this, we adopt per-pixel Plücker ray injection alongside resolution-aware 3D RoPE in self-attention, building both camera conditioning and spatial positional encoding on the target latent to jointly control camera position, focal length, and native resolution editing without cropping or outpainting. To overcome the scarcity of paired training data, we further develop a scalable synthetic pipeline that constructs diverse dynamic scenes through structured multi-camera recording and generates synchronized videos with varied camera configurations. With a tailored orthogonal training strategy, CameraAnything enables expressive video reshooting with arbitrary viewpoint control, focal length adjustment, resolution adaptation, and multi-shot transitions within a single generation process, offering strong practical value for cinematic video editing and cross-platform content adaptation in video production.