English
Related papers

Related papers: OpenVE-3M: A Large-Scale High-Quality Dataset for …

200 papers

Recent advancements of generative AI have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. In doing so, these studies mainly transfer the inherent motion…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Chang Liu , Rui Li , Kaidong Zhang , Yunwei Lan , Dong Liu

We propose MLV-Edit, a training-free, flow-based framework that address the unique challenges of minute-level video editing. While existing techniques excel in short-form video manipulation, scaling them to long-duration videos remains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Yangyi Cao , Yuanhang Li , Lan Chen , Qi Mao

The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm comparisons difficult, while video editing lacks reliable…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Lifan Jiang , Tianrun Wu , Yuhang Pei , Chenyang Wang , Boxi Wu , Deng Cai

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Jiafeng Liang , Shixin Jiang , Zekun Wang , Haojie Pan , Zerui Chen , Zheng Chu , Ming Liu , Ruiji Fu , Zhongyuan Wang , Bing Qin

Recent advances in text-to-image (T2I) diffusion models have significantly improved semantic image editing, yet most methods fall short in performing 3D-aware object manipulation. In this work, we present FFSE, a 3D-aware autoregressive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xincheng Shuai , Zhenyuan Qin , Henghui Ding , Dacheng Tao

Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity of inter-person interactions. We introduce the task of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yebin Yang , Di Wen , Lei Qi , Weitong Kong , Junwei Zheng , Ruiping Liu , Yufan Chen , Chengzhi Wu , Kailun Yang , Yuqian Fu , Danda Pani Paudel , Luc Van Gool , Kunyu Peng

We consider the problem of editing 3D objects and scenes based on open-ended language instructions. A common approach to this problem is to use a 2D image generator or editor to guide the 3D editing process, obviating the need for 3D data.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Minghao Chen , Iro Laina , Andrea Vedaldi

There are substantial instructional videos on the Internet, which enables us to acquire knowledge for completing various tasks. However, most existing datasets for instructional video analysis have the limitations in diversity and…

Computer Vision and Pattern Recognition · Computer Science 2019-03-08 Yansong Tang , Dajun Ding , Yongming Rao , Yu Zheng , Danyang Zhang , Lili Zhao , Jiwen Lu , Jie Zhou

Recent large vision-language models (LVLMs) for video understanding are primarily fine-tuned with various videos scraped from online platforms. Existing datasets, such as ActivityNet, require considerable human labor for structuring and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Zhende Song , Chenchen Wang , Jiamu Sheng , Chi Zhang , Shengji Tang , Jiayuan Fan , Tao Chen

High-quality and open datasets remain a major bottleneck for text-to-image (T2I) fine-tuning. Despite rapid progress in model architectures and training pipelines, most publicly available fine-tuning datasets suffer from low resolution,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Xu Ma , Yitian Zhang , Qihua Dong , Yun Fu

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Qianqian Qiao , DanDan Zheng , Yihang Bo , Bao Peng , Heng Huang , Longteng Jiang , Huaye Wang , Jingdong Chen , Jun Zhou , Xin Jin

Large Language Models (LLMs) have significantly advanced natural language processing, demonstrating strong capabilities in tasks such as text generation, summarization, and reasoning. Recently, their potential for automating precise text…

Computation and Language · Computer Science 2026-01-27 Yiming Zeng , Wanhao Yu , Zexin Li , Tao Ren , Yu Ma , Jinghan Cao , Xiyan Chen , Tingting Yu

High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human…

Computation and Language · Computer Science 2024-10-08 Zhangchen Xu , Fengqing Jiang , Luyao Niu , Yuntian Deng , Radha Poovendran , Yejin Choi , Bill Yuchen Lin

While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical failure mode crucial in professional applications: the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yujia Yang , Yuanxiang Wang , Zhenyu Guan , Tiankun Yang , Chenxi Bao , Haopeng Jin , Jinwen Luo , Xinyu Zuo , Lisheng Duan , Haijin Liang , Jin Ma , Xinming Wang , Ruiwen Tao , Hongzhu Yi

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models present impressive…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Bo Zhao , Boya Wu , Muyang He , Tiejun Huang

We present Interactive Neural Video Editing (INVE), a real-time video editing solution, which can assist the video editing process by consistently propagating sparse frame edits to the entire video clip. Our method is inspired by the recent…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Jiahui Huang , Leonid Sigal , Kwang Moo Yi , Oliver Wang , Joon-Young Lee

Instruction-based image editing (IIE) aims to modify images according to textual instructions while preserving irrelevant content. Despite recent advances in diffusion transformers, existing methods often suffer from over-editing,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Jingxuan He , Xiyu Wang , Mengyu Zheng , Xiangyu Zeng , Yunke Wang , Chang Xu

We introduce OpenShape, a method for learning multi-modal joint representations of text, image, and point clouds. We adopt the commonly used multi-modal contrastive learning framework for representation alignment, but with a specific focus…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Minghua Liu , Ruoxi Shi , Kaiming Kuang , Yinhao Zhu , Xuanlin Li , Shizhong Han , Hong Cai , Fatih Porikli , Hao Su

The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Purui Bai , Tao Wu , Jiayang Sun , Xinyue Liu , Huaibo Huang , Ran He

Although existing video editing methods are generally feasible, they often require many costly iterations and still struggle to deliver high-quality yet satisfying editing results. We attribute this limitation to the prevalent data-to-data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Guanlong Jiao , Chenyangguang Zhang , Jia Jun Cheng Xian , Zewei Zhang , Renjie Liao