Generative Video Propagation
Shaoteng Liu, Tianyu Wang, Jui-Hsien Wang, Qing Liu, Zhifei Zhang, Joon-Young Lee, Yijun Li, Bei Yu, Zhe Lin, Soo Ye Kim, Jiaya Jia
CVPR · 2025 · Tech Lead
Soo Ye Kim: Co-corresponding author; Jiaya Jia: Co-corresponding author
Research summary
A unified framework that propagates edits from the first frame to all following frames using video generation models.
Abstract
Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework.
Key insight
Repurposes video generation models to propagate first-frame edits throughout an entire video, enabling diverse manipulation tasks including editing, object insertion, removal, and tracking in a single framework.
Technical connections & research relevance
Technical themes
- first-frame edit propagation
- temporally coherent video editing
- visual domain augmentation
- embodied AI research direction
- sim2real transfer research direction
Method connections
GenProp propagates first-frame edits through video while preserving surrounding content, enabling changes to objects, backgrounds, shadows, and reflections.
Research relevance
These capabilities suggest a route to temporally coherent visual augmentation for embodied AI: varying the appearance of recorded or rendered observations to study perception robustness and sim2real transfer.
Evaluation scope
Robotics performance remains an open evaluation direction. The paper evaluates video manipulation rather than robot policy transfer; the model produces edited video, not an executable physics simulator.
Primary sources
CVPR 2025 proceedings; arXiv preprint also available
Paper & resources
BibTeX
@inproceedings{liu2025genprop,
title={Generative Video Propagation},
author={Liu, Shaoteng and Wang, Tianyu and Wang, Jui-Hsien and Liu, Qing and Zhang, Zhifei and Lee, Joon-Young and Li, Yijun and Yu, Bei and Lin, Zhe and Kim, Soo Ye and Jia, Jiaya},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
pages={17712--17722},
year={2025}
}