English
Related papers

Related papers: ESA: Energy-Based Shot Assembly Optimization for A…

200 papers

We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Learned Clip Assembly (LCA) score, a learning-based metric that…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Chen Yi Lu , Md Mehrab Tanjim , Ishita Dasgupta , Somdeb Sarkhel , Gang Wu , Saayan Mitra , Somali Chaterji

Video summarization methods are usually classified into shot-level or frame-level methods, which are individually used in a general way. This paper investigates the underlying complementarity between the frame-level and shot-level methods,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Yubo An , Shenghui Zhao , Guoqiang Zhang

Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e.g., VLM features or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Xinyao Zhang , Wenkai Dong , Yuxin Song , Bo Fang , Qi Zhang , Jing Wang , Fan Chen , Hui Zhang , Haocheng Feng , Yu Lu , Hang Zhou , Chun Yuan , Jingdong Wang

Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-world complexity. Pre-trained vision-language models tend to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Muhammad Huzaifa , Yova Kementchedjhieva

Automatic video editing involving at least the steps of selecting the most valuable footage from points of view of visual quality and the importance of action filmed; and cutting the footage into a brief and coherent visual story that would…

Computer Vision and Pattern Recognition · Computer Science 2019-07-18 Sergey Podlesnyy

The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Alejandro Pardo , Jui-Hsien Wang , Bernard Ghanem , Josef Sivic , Bryan Russell , Fabian Caba Heilbron

Language is compositional; an instruction can express multiple relation constraints to hold among objects in a scene that a robot is tasked to rearrange. Our focus in this work is an instructable scene-rearranging framework that generalizes…

Motivated by the superior performance of image diffusion models, more and more researchers strive to extend these models to the text-based video editing task. Nevertheless, current video editing tasks mainly suffer from the dilemma between…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Yutao Chen , Xingning Dong , Tian Gan , Chunluan Zhou , Ming Yang , Qingpei Guo

We propose EMMA, an efficient and unified architecture for multimodal understanding, generation and editing. Specifically, EMMA primarily consists of 1) An efficient autoencoder with a 32x compression ratio, which significantly reduces the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xin He , Longhui Wei , Jianbo Ouyang , Minghui Liao , Lingxi Xie , Qi Tian

Active learning enhances annotation efficiency by selecting the most revealing samples for labeling, thereby reducing reliance on extensive human input. Previous methods in semantic segmentation have centered on individual pixels or small…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jinchao Ge , Zeyu Zhang , Minh Hieu Phan , Bowen Zhang , Akide Liu , Yang Zhao , Shuwen Zhao

The possibility of sharing one's point of view makes use of wearable cameras compelling. These videos are often long, boring and coupled with extreme shake, as the camera is worn on a moving person. Fast forwarding (i.e. frame sampling) is…

Computer Vision and Pattern Recognition · Computer Science 2017-01-13 Tavi Halperin , Yair Poleg , Chetan Arora , Shmuel Peleg

We present GAZED- eye GAZe-guided EDiting for videos captured by a solitary, static, wide-angle and high-resolution camera. Eye-gaze has been effectively employed in computational applications as a cue to capture interesting scene content;…

Computer Vision and Pattern Recognition · Computer Science 2020-10-23 K L Bhanu Moorthy , Moneish Kumar , Ramanathan Subramaniam , Vineet Gandhi

Perspective texture synthesis has great significance in many fields like video editing, scene capturing etc., due to its ability to read and control global feature information. In this paper, we present a novel example-based, specifically…

Computer Vision and Pattern Recognition · Computer Science 2020-06-28 Syed Muhammad Arsalan Bashir , Farhan Ali Khan Ghouri

Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent editing of subjects and scenes remains a critical challenge. Recent attempts to incorporate richer…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Fengyuan Yang , Luying Huang , Jiazhi Guan , Quanwei Yang , Dongwei Pan , Jianglin Fu , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou , Angela Yao

In few-shot learning, the selection of samples has a significant impact on the performance of the model. While effective sample selection strategies are well-established in supervised settings, research on large language models largely…

Machine Learning · Computer Science 2026-04-20 Branislav Pecher , Ivan Srba , Maria Bielikova , Joaquin Vanschoren

Current diffusion-based video editing primarily focuses on local editing (\textit{e.g.,} object/background editing) or global style editing by utilizing various dense correspondences. However, these methods often fail to accurately edit the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Xiangpeng Yang , Linchao Zhu , Hehe Fan , Yi Yang

Current video-language models (VLMs) rely extensively on instance-level alignment between video and language modalities, which presents two major limitations: (1) visual reasoning disobeys the natural perception that humans do in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-04 Khoa Vo , Thinh Phan , Kashu Yamazaki , Minh Tran , Ngan Le

We propose a zero-shot approach to image harmonization, aiming to overcome the reliance on large amounts of synthetic composite images in existing methods. These methods, while showing promising results, involve significant training…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Jianqi Chen , Yilan Zhang , Zhengxia Zou , Keyan Chen , Zhenwei Shi

The goal in episodic memory (EM) is to search a long egocentric video to answer a natural language query (e.g., "where did I leave my purse?"). Existing EM methods exhaustively extract expensive fixed-length clip features to look everywhere…

Computer Vision and Pattern Recognition · Computer Science 2023-06-29 Santhosh Kumar Ramakrishnan , Ziad Al-Halah , Kristen Grauman

Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Youyuan Zhang , Xuan Ju , James J. Clark
‹ Prev 1 2 3 10 Next ›