English
Related papers

Related papers: MF2Summ: Multimodal Fusion for Video Summarization…

200 papers

To accomplish punctuation restoration, most existing methods focus on introducing extra information (e.g., part-of-speech) or addressing the class imbalance problem. Recently, large-scale transformer-based pre-trained language models (PLMS)…

Computation and Language · Computer Science 2022-11-10 Yangjun Wu , Kebin Fang , Yao Zhao , Hao Zhang , Lifeng Shi , Mengqi Zhang

Video smmarization is a crucial method to reduce the time of videos which reduces the spent time to watch/review a long video. This apporach has became more important as the amount of publisehed video is increasing everyday. A single or…

Computer Vision and Pattern Recognition · Computer Science 2024-01-22 Vahid Ahmadi Kalkhorani , Qingquan Zhang , Guanqun Song , Ting Zhu

Current video summarization methods rely heavily on supervised computer vision techniques, which demands time-consuming and subjective manual annotations. To overcome these limitations, we investigated self-supervised video summarization.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Tomoya Sugihara , Shuntaro Masuda , Ling Xiao , Toshihiko Yamasaki

The video topic segmentation (VTS) task segments videos into intelligible, non-overlapping topics, facilitating efficient comprehension of video content and quick access to specific content. VTS is also critical to various downstream video…

Artificial Intelligence · Computer Science 2024-12-31 Hai Yu , Chong Deng , Qinglin Zhang , Jiaqing Liu , Qian Chen , Wen Wang

For multimodal tasks, a good feature extraction network should extract information as much as possible and ensure that the extracted feature embedding and other modal feature embedding have an excellent mutual understanding. The latter is…

Computer Vision and Pattern Recognition · Computer Science 2021-06-01 Jianning Wu , Zhuqing Jiang , Shiping Wen , Aidong Men , Haiying Wang

For many applications with limited computation, communication, storage and energy resources, there is an imperative need of computer vision methods that could select an informative subset of the input video for efficient processing at or…

Computer Vision and Pattern Recognition · Computer Science 2018-05-09 Shuyue Lan , Rameswar Panda , Qi Zhu , Amit K. Roy-Chowdhury

The rapid rise of video content on platforms such as TikTok and YouTube has transformed information dissemination, but it has also facilitated the spread of harmful content, particularly hate videos. Despite significant efforts to combat…

Multimedia · Computer Science 2025-05-20 Yinghui Zhang , Tailin Chen , Yuchen Zhang , Zeyu Fu

Recently, the research interest of person re-identification (ReID) has gradually turned to video-based methods, which acquire a person representation by aggregating frame features of an entire video. However, existing video-based ReID…

Computer Vision and Pattern Recognition · Computer Science 2020-09-14 Xinyang Jiang , Yifei Gong , Xiaowei Guo , Qize Yang , Feiyue Huang , Weishi Zheng , Feng Zheng , Xing Sun

Few-shot 3D point cloud segmentation (FS-PCS) aims at generalizing models to segment novel categories with minimal annotated support samples. While existing FS-PCS methods have shown promise, they primarily focus on unimodal point cloud…

Computer Vision and Pattern Recognition · Computer Science 2025-02-27 Zhaochong An , Guolei Sun , Yun Liu , Runjia Li , Min Wu , Ming-Ming Cheng , Ender Konukoglu , Serge Belongie

Recognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehension. Most previous…

Computer Vision and Pattern Recognition · Computer Science 2022-07-13 Jiashuo Yu , Ying Cheng , Rui-Wei Zhao , Rui Feng , Yuejie Zhang

Exploiting both audio and visual modalities for video classification is a challenging task, as the existing methods require large model architectures, leading to high computational complexity and resource requirements. Smaller…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Mahrukh Awan , Asmar Nadeem , Muhammad Junaid Awan , Armin Mustafa , Syed Sameed Husain

We present Attend-Fusion, a novel and efficient approach for audio-visual fusion in video classification tasks. Our method addresses the challenge of exploiting both audio and visual modalities while maintaining a compact model…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Mahrukh Awan , Asmar Nadeem , Armin Mustafa

Large Multimodal Models (LMMs) for video-audio understanding have traditionally been evaluated only on shorter videos of a few minutes long. In this paper, we introduce QMAVIS (Q Team-Multimodal Audio Video Intelligent Sensemaking), a novel…

Artificial Intelligence · Computer Science 2026-01-13 Zixing Lin , Jiale Wang , Gee Wah Ng , Lee Onn Mak , Chan Zhi Yang Jeriel , Jun Yang Lee , Yaohao Li

Multimodal information (e.g., visual, acoustic, and textual) has been widely used to enhance representation learning for micro-video recommendation. For integrating multimodal information into a joint representation of micro-video,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Han Liu , Yinwei Wei , Fan Liu , Wenjie Wang , Liqiang Nie , Tat-Seng Chua

Multimodal summarization usually suffers from the problem that the contribution of the visual modality is unclear. Existing multimodal summarization approaches focus on designing the fusion methods of different modalities, while ignoring…

Computation and Language · Computer Science 2023-07-07 Min Xiao , Junnan Zhu , Haitao Lin , Yu Zhou , Chengqing Zong

Transitioning Multimodal Large Language Models (MLLMs) from offline to online streaming video understanding is essential for continuous perception. However, existing methods lack flexible adaptivity, leading to irreversible detail loss and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Kangcong Li , Peng Ye , Lin Zhang , Chao Wang , Huafeng Qin , Tao Chen

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable…

Computer Vision and Pattern Recognition · Computer Science 2016-11-11 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

In this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and smartphones. The…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Ioannis Kontostathis , Evlampios Apostolidis , Vasileios Mezaris

Recently, with the emergence of large language models, multimodal LLMs have demonstrated exceptional capabilities in image and video modalities. Despite advancements in video comprehension, the substantial computational demands of long…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Ming Nie , Chunwei Wang , Hang Xu , Li Zhang

This paper introduces a novel variant of video summarization, namely building a summary that depends on the particular aspect of a video the viewer focuses on. We refer to this as $\textit{viewpoint}$. To infer what the desired…

Computer Vision and Pattern Recognition · Computer Science 2018-04-11 Atsushi Kanehira , Luc Van Gool , Yoshitaka Ushiku , Tatsuya Harada