English
Related papers

Related papers: Uncertainty-Aware and Decoder-Aligned Learning for…

200 papers

Video summarization aims at generating concise video summaries from the lengthy videos, to achieve better user watching experience. Due to the subjectivity, purely supervised methods for video summarization may bring the inherent errors…

Computer Vision and Pattern Recognition · Computer Science 2020-07-30 Tianyu Liu

Long video understanding is heavily bottlenecked by a rigid one-shot paradigm: existing methods either densely encode videos at prohibitive memory and latency costs, or aggressively compress them into sparse frame sets that irreversibly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Xiao Yang , Yingzhe Ma , Haoxuan Yu , Zixin Li , Ning Qin

Selecting informative keyframes is critical for efficient video understanding, yet existing approaches often rely on heuristics, ignore semantics, or produce redundant frames. We propose KeyScore, a caption-aware frame scoring method that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Shih-Yao Lin , Sibendu Paul , Caren Chen

Video summarization is an effective way to facilitate video searching and browsing. Most of existing systems employ encoder-decoder based recurrent neural networks, which fail to explicitly diversify the system-generated summary frames…

Computer Vision and Pattern Recognition · Computer Science 2020-09-24 Ping Li , Qinghao Ye , Luming Zhang , Li Yuan , Xianghua Xu , Ling Shao

Existing video copy detection methods generally measure video similarity based on spatial similarities between key frames, neglecting the latent similarity in temporal dimension, so that the video similarity is biased towards spatial…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Zhen Han , Xiangteng He , Mingqian Tang , Yiliang Lv

We propose a rubric-guided, pseudo-labeled, and prompt-driven zero-shot video summarization framework that bridges large language models with structured semantic reasoning. A small subset of human annotations is converted into…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Yuanli Wu , Long Zhang , Yue Du , Bin Li

Video captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence. The encoder-decoder framework is the most popular paradigm for this task in recent years. However, there exist some…

Computer Vision and Pattern Recognition · Computer Science 2021-02-15 Haoran Chen , Jianmin Li , Xiaolin Hu

Audio-Visual Speech Recognition (AVSR) seeks to model, and thereby exploit, the dynamic relationship between a human voice and the corresponding mouth movements. A recently proposed multimodal fusion strategy, AV Align, based on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-20 George Sterpu , Christian Saam , Naomi Harte

This paper proposes an efficient video summarization framework that will give a gist of the entire video in a few key-frames or video skims. Existing video summarization frameworks are based on algorithms that utilize computer vision…

Computer Vision and Pattern Recognition · Computer Science 2021-01-28 Sai Sukruth Bezugam , Swatilekha Majumdar , Chetan Ralekar , Tapan Kumar Gandhi

Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when only a small number of frames can be retained, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Mengyu Zhao , Di Fu , Yongyu Xie , Jiaxing Zhang , Zhigang Yuan , Shirin Jalali , Yong Cao

We address the problem of highlight detection from a 360 degree video by summarizing it both spatially and temporally. Given a long 360 degree video, we spatially select pleasantly-looking normal field-of-view (NFOV) segments from unlimited…

Computer Vision and Pattern Recognition · Computer Science 2018-02-01 Youngjae Yu , Sangho Lee , Joonil Na , Jaeyun Kang , Gunhee Kim

Deep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation between successive frames and then compress the residual…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Wufei Ma , Jiahao Li , Bin Li , Yan Lu

We present a self-supervised approach for learning video representations using temporal video alignment as a pretext task, while exploiting both frame-level and video-level information. We leverage a novel combination of temporal alignment…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Sanjay Haresh , Sateesh Kumar , Huseyin Coskun , Shahram Najam Syed , Andrey Konin , Muhammad Zeeshan Zia , Quoc-Huy Tran

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Bo Fang , Wenhao Wu , Chang Liu , Yu Zhou , Yuxin Song , Weiping Wang , Xiangbo Shu , Xiangyang Ji , Jingdong Wang

The recent development of Video-based Large Language Models (VideoLLMs), has significantly advanced video summarization by aligning video features and, in some cases, audio features with Large Language Models (LLMs). Each of these VideoLLMs…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Kuan-Chen Mu , Zhi-Yi Chin , Wei-Chen Chiu

Recent years have witnessed a resurgence of interest in video summarization. However, one of the main obstacles to the research on video summarization is the user subjectivity - users have various preferences over the summaries. The…

Computer Vision and Pattern Recognition · Computer Science 2017-07-18 Aidean Sharghi , Jacob S. Laurel , Boqing Gong

Inspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder framework incorporating multimodal semantic attributes for…

Computer Vision and Pattern Recognition · Computer Science 2019-05-09 Liang Sun , Bing Li , Chunfeng Yuan , Zhengjun Zha , Weiming Hu

This paper addresses automatic summarization and search in visual data comprising of videos, live streams and image collections in a unified manner. In particular, we propose a framework for multi-faceted summarization which extracts…

Computer Vision and Pattern Recognition · Computer Science 2017-04-06 Anurag Sahoo , Vishal Kaushal , Khoshrav Doctor , Suyash Shetty , Rishabh Iyer , Ganesh Ramakrishnan

The topic diversity of open-domain videos leads to various vocabularies and linguistic expressions in describing video contents, and therefore, makes the video captioning task even more challenging. In this paper, we propose an unified…

Computer Vision and Pattern Recognition · Computer Science 2023-02-15 Shizhe Chen , Jia Chen , Qin Jin , Alexander Hauptmann

While multi-modal learning has advanced significantly, current approaches often create inconsistencies in representation and reasoning of different modalities. We propose UMaT, a theoretically-grounded framework that unifies visual and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Xiaowei Bi , Zheyuan Xu
‹ Prev 1 4 5 6 7 8 10 Next ›