{"componentChunkName":"component---src-templates-publication-js","path":"/research/genprop-cvpr2025/","result":{"pageContext":{"publication":{"id":"genprop-cvpr2025","title":"Generative Video Propagation","authors":["Shaoteng Liu","Tianyu Wang","Jui-Hsien Wang","Qing Liu","Zhifei Zhang","Joon-Young Lee","Yijun Li","Bei Yu","Zhe Lin","Soo Ye Kim","Jiaya Jia"],"highlightAuthor":"Tianyu Wang","venue":"CVPR","year":2025,"tldr":"A unified framework that propagates edits from the first frame to all following frames using video generation models.","links":{"paper":"https://openaccess.thecvf.com/content/CVPR2025/html/Liu_Generative_Video_Propagation_CVPR_2025_paper.html","project":"https://genprop.github.io/","preprint":"https://arxiv.org/abs/2412.19761"},"insight":"Repurposes video generation models to propagate first-frame edits throughout an entire video, enabling diverse manipulation tasks including editing, object insertion, removal, and tracking in a single framework.","teaser":"/images/information/genprop.png","teaserVideo":"/videos/genprop-cvpr2025.mp4","teaserGif":"/videos/previews/genprop-cvpr2025.gif","bibtex":"@inproceedings{liu2025genprop,\n  title={Generative Video Propagation},\n  author={Liu, Shaoteng and Wang, Tianyu and Wang, Jui-Hsien and Liu, Qing and Zhang, Zhifei and Lee, Joon-Young and Li, Yijun and Yu, Bei and Lin, Zhe and Kim, Soo Ye and Jia, Jiaya},\n  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},\n  pages={17712--17722},\n  year={2025}\n}","role":"Tech Lead","authorMarkers":{"Soo Ye Kim":["†"],"Jiaya Jia":["†"]},"markerLegend":{"†":"Co-corresponding author"},"abstract":"Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of such models. Specifically, our framework, GenProp, encodes the original video with a selective content encoder and propagates the changes made to the first frame using an image-to-video generation model. We propose a data generation scheme to cover multiple video tasks based on instance-level video segmentation datasets. Our model is trained by incorporating a mask prediction decoder head and optimizing a region-aware loss to aid the encoder to preserve the original content while the generation model propagates the modified region. This novel design opens up new possibilities: In editing scenarios, GenProp allows substantial changes to an object's shape; for insertion, the inserted objects can exhibit independent motion; for removal, GenProp effectively removes effects like shadows and reflections from the whole video; for tracking, GenProp is capable of tracking objects and their associated effects together. Experiment results demonstrate the leading performance of our model in various video tasks, and we further provide in-depth analyses of the proposed framework.","abstractSource":"https://openaccess.thecvf.com/content/CVPR2025/html/Liu_Generative_Video_Propagation_CVPR_2025_paper.html","version":"CVPR 2025 proceedings; arXiv preprint also available","technicalContext":{"themes":["first-frame edit propagation","temporally coherent video editing","visual domain augmentation","embodied AI research direction","sim2real transfer research direction"],"method":"GenProp propagates first-frame edits through video while preserving surrounding content, enabling changes to objects, backgrounds, shadows, and reflections.","relevance":"These capabilities suggest a route to temporally coherent visual augmentation for embodied AI: varying the appearance of recorded or rendered observations to study perception robustness and sim2real transfer.","scope":"Robotics performance remains an open evaluation direction. The paper evaluates video manipulation rather than robot policy transfer; the model produces edited video, not an executable physics simulator.","sources":[{"label":"GenProp: Sections 3 and 4.2, supplementary Section S6","url":"https://arxiv.org/html/2412.19761v1"}],"description":"GenProp propagates first-frame edits through video. Its temporally coherent visual augmentation suggests embodied AI and sim2real research directions requiring robotics validation."}}}},"staticQueryHashes":["63159454"]}