English
Related papers

Related papers: Learning to Unify Audio, Visual and Text for Audio…

200 papers

Open-Vocabulary Temporal Action Localization (OVTAL) enables a model to recognize any desired action category in videos without the need to explicitly curate training data for all categories. However, this flexibility poses significant…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Akshita Gupta , Aditya Arora , Sanath Narayan , Salman Khan , Fahad Shahbaz Khan , Graham W. Taylor

In this paper, we introduce a new problem, named audio-visual video parsing, which aims to parse a video into temporal event segments and label them as either audible, visible, or both. Such a problem is essential for a complete…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Yapeng Tian , Dingzeyu Li , Chenliang Xu

Audio-Visual Source Localization (AVSL) aims to locate sounding objects within video frames given the paired audio clips. Existing methods predominantly rely on self-supervised contrastive learning of audio-visual correspondence. Without…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Yuxin Guo , Shijie Ma , Hu Su , Zhiqing Wang , Yuhao Zhao , Wei Zou , Siyang Sun , Yun Zheng

Temporal Action Localization (TAL) requires identifying both the boundaries and categories of actions in untrimmed videos. While vision-language models (VLMs) offer rich semantics to complement visual evidence, existing approaches tend to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jiaqi Li , Guangming Wang , Shuntian Zheng , Minzhe Ni , Xiaoman Lu , Guanghui Ye , Yu Guan

Text-video retrieval is a challenging task that aims to search relevant video contents based on natural language descriptions. The key to this problem is to measure text-video similarities in a joint embedding space. However, most existing…

Computer Vision and Pattern Recognition · Computer Science 2021-04-21 Xiaohan Wang , Linchao Zhu , Yi Yang

Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Wentao Zhu

Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Yung-Hsuan Lai , Janek Ebbers , Yu-Chiang Frank Wang , François Germain , Michael Jeffrey Jones , Moitreya Chatterjee

We present a simplified, task-agnostic multi-modal pre-training approach that can accept either video or text input, or both for a variety of end tasks. Existing pre-training are task-specific by adopting either a single cross-modal encoder…

Computer Vision and Pattern Recognition · Computer Science 2021-10-04 Hu Xu , Gargi Ghosh , Po-Yao Huang , Prahal Arora , Masoumeh Aminzadeh , Christoph Feichtenhofer , Florian Metze , Luke Zettlemoyer

Vision-language alignment learning for video-text retrieval arouses a lot of attention in recent years. Most of the existing methods either transfer the knowledge of image-text pretraining model to video-text retrieval task without fully…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Yizhen Chen , Jie Wang , Lijian Lin , Zhongang Qi , Jin Ma , Ying Shan

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Sarah Ibrahimi , Xiaohang Sun , Pichao Wang , Amanmeet Garg , Ashutosh Sanan , Mohamed Omar

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip movements, VSR tasks require effectively utilizing any…

Sound · Computer Science 2024-10-23 Zehua Liu , Xiaolou Li , Chen Chen , Li Guo , Lantian Li , Dong Wang

Temporal Action Localization (TAL) aims to identify actions' start, end, and class labels in untrimmed videos. While recent advancements using transformer networks and Feature Pyramid Networks (FPN) have enhanced visual feature recognition…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Edward Fish , Jon Weinbren , Andrew Gilbert

Visual and audio events simultaneously occur and both attract attention. However, most existing saliency prediction works ignore the influence of audio and only consider vision modality. In this paper, we propose a multitask learning method…

Computer Vision and Pattern Recognition · Computer Science 2021-11-17 Minglang Qiao , Yufan Liu , Mai Xu , Xin Deng , Bing Li , Weiming Hu , Ali Borji

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize…

Sound · Computer Science 2025-08-07 Yu Chen , Hongxu Zhu , Jiadong Wang , Kainan Chen , Xinyuan Qian

Although several research works have been reported on audio-visual sound source localization in unconstrained videos, no datasets and metrics have been proposed in the literature to quantitatively evaluate its performance. Defining the…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Tokuhiro Nishikawa , Daiki Shimada , Jerry Jun Yokono

With the recent success of the pre-training technique for NLP and image-linguistic tasks, some video-linguistic pre-training works are gradually developed to improve video-text related downstream tasks. However, most of the existing…

Computer Vision and Pattern Recognition · Computer Science 2020-09-16 Huaishao Luo , Lei Ji , Botian Shi , Haoyang Huang , Nan Duan , Tianrui Li , Jason Li , Taroon Bharti , Ming Zhou

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represent knowledge using…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-22 Xianghu Yue , Xiaohai Tian , Lu Lu , Malu Zhang , Zhizheng Wu , Haizhou Li

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Bo Fang , Wenhao Wu , Chang Liu , Yu Zhou , Yuxin Song , Weiping Wang , Xiangbo Shu , Xiangyang Ji , Jingdong Wang

We present an audio-visual multimodal approach for the task of zeroshot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent past but has primarily been limited to visual modality and to…

Computer Vision and Pattern Recognition · Computer Science 2019-10-22 Kranti Kumar Parida , Neeraj Matiyali , Tanaya Guha , Gaurav Sharma

Weakly supervised temporal action localization (WTAL) aims to detect action instances in untrimmed videos using only video-level annotations. Since many existing works optimize WTAL models based on action classification labels, they…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Geuntaek Lim , Hyunwoo Kim , Joonsoo Kim , Yukyung Choi