English
Related papers

Related papers: Beyond Generation: Unlocking Universal Editing via…

200 papers

Video Instance Segmentation (VIS) faces significant annotation challenges due to its dual requirements of pixel-level masks and temporal consistency labels. While recent unsupervised methods like VideoCutLER eliminate optical flow…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Kaixuan Lu , Mehmet Onurcan Kaya , Dim P. Papadopoulos

Recent advances in large generative models have greatly enhanced both image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Jay Zhangjie Wu , Xuanchi Ren , Tianchang Shen , Tianshi Cao , Kai He , Yifan Lu , Ruiyuan Gao , Enze Xie , Shiyi Lan , Jose M. Alvarez , Jun Gao , Sanja Fidler , Zian Wang , Huan Ling

Current text-driven image editing methods typically follow one of two directions: relying on large-scale, high-quality editing pair datasets to improve editing precision and diversity, or exploring alternative dataset-free techniques.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Chenrui Ma , Xi Xiao , Tianyang Wang , Yanning Shen

Ultrasound (US) is widely used for its advantages of real-time imaging, radiation-free and portability. In clinical practice, analysis and diagnosis often rely on US sequences rather than a single image to obtain dynamic anatomical…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Jiamin Liang , Xin Yang , Yuhao Huang , Kai Liu , Xinrui Zhou , Xindi Hu , Zehui Lin , Huanjia Luo , Yuanji Zhang , Yi Xiong , Dong Ni

In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Panwen Hu , Nan Xiao , Feifei Li , Yongquan Chen , Rui Huang

Recent progress in controllable image generation and editing is largely driven by diffusion-based methods. Although diffusion models perform exceptionally well in specific tasks with tailored designs, establishing a unified model is still…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Jiteng Mu , Nuno Vasconcelos , Xiaolong Wang

The In-context generation paradigm recently has demonstrated strong power in instructional image editing with both data efficiency and synthesis quality. Nevertheless, shaping such in-context learning for instruction-based video editing is…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Zhongwei Zhang , Fuchen Long , Wei Li , Zhaofan Qiu , Wu Liu , Ting Yao , Tao Mei

AI-generated face detectors trained via supervised learning typically rely on synthesized images from specific generators, limiting their generalization to emerging generative techniques. To overcome this limitation, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Mian Zou , Nan Zhong , Baosheng Yu , Yibing Zhan , Kede Ma

Self-supervised learning has emerged as a powerful paradigm for label-free model pretraining, particularly in the video domain, where manual annotation is costly and time-intensive. However, existing self-supervised approaches employ…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Akash Kumar , Ashlesha Kumar , Vibhav Vineet , Yogesh S Rawat

Instruction-guided 3D editing is a rapidly emerging field with the potential to broaden access to 3D content creation. However, existing methods face critical limitations: optimization-based approaches are prohibitively slow, while…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Weiwei Cai , Shuangkang Fang , Weicai Ye , Xin Dong , Yunhan Yang , Xuanyang Zhang , Wei Cheng , Yanpei Cao , Gang Yu , Tao Chen

Generative AI presents transformative potential across various domains, from creative arts to scientific visualization. However, the utility of AI-generated imagery is often compromised by visual flaws, including anatomical inaccuracies,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Zhenyu Yu , Chee Seng Chan

Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Ming Li , Xin Gu , Fan Chen , Xiaoying Xing , Longyin Wen , Chen Chen , Sijie Zhu

Video content creation keeps growing at an incredible pace; yet, creating engaging stories remains challenging and requires non-trivial video editing expertise. Many video editing components are astonishingly hard to automate primarily due…

Computer Vision and Pattern Recognition · Computer Science 2021-09-30 Alejandro Pardo , Fabian Caba Heilbron , Juan León Alcázar , Ali Thabet , Bernard Ghanem

Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as…

Artificial Intelligence · Computer Science 2026-02-13 Lingyong Yan , Jiulong Wu , Dong Xie , Weixian Shi , Deguo Xia , Jizhou Huang

Video generation and editing conditioned on text prompts or images have undergone significant advancements. However, challenges remain in accurately controlling global layout and geometry details solely by texts, and supporting motion…

Graphics · Computer Science 2025-04-01 Feng-Lin Liu , Hongbo Fu , Xintao Wang , Weicai Ye , Pengfei Wan , Di Zhang , Lin Gao

Video generation models have significantly advanced embodied intelligence, unlocking new possibilities for generating diverse robot data that capture perception, reasoning, and action in the physical world. However, synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Yufan Deng , Zilin Pan , Hongyu Zhang , Xiaojie Li , Ruoqing Hu , Yufei Ding , Yiming Zou , Yan Zeng , Daquan Zhou

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Shangkun Sun , Xiaoyu Liang , Songlin Fan , Wenxu Gao , Wei Gao

Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to align with complex text prompts, particularly when multiple objects, attributes, or spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Daeun Lee , Jaehong Yoon , Jaemin Cho , Mohit Bansal

While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance under complex interleaved instructions. This limitation stems…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Yabo Zhang , Kunchang Li , Dewei Zhou , Xinyu Huang , Xun Wang

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entangle these factors, limiting independent control. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Yukun Wang , Ruihuang Li , Jiale Tao , Shiyuan Yang , Liyi Chen , Zhantao Yang , Handz , Yulan Guo , Shuai Shao , Qinglin Lu