English
Related papers

Related papers: TEn-CATG:Text-Enriched Audio-Visual Video Parsing …

200 papers

Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most state-of-the-art TVR methods learn image-to-video transfer learning based on large-scale pre-trained visionlanguage models (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Meng Cao , Haoran Tang , Jinfa Huang , Peng Jin , Can Zhang , Ruyang Liu , Long Chen , Xiaodan Liang , Li Yuan , Ge Li

The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explored still have…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Xiangyu Wu , Qing-Yuan Jiang , Yang Yang , Yi-Feng Wu , Qing-Guo Chen , Jianfeng Lu

Audio-visual segmentation (AVS) aims to segment sound sources in the video sequence, requiring a pixel-level understanding of audio-visual correspondence. As the Segment Anything Model (SAM) has strongly impacted extensive fields of dense…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Juhyeong Seon , Woobin Im , Sebin Lee , Jumin Lee , Sung-Eui Yoon

Temporal Video Grounding (TVG), which requires pinpointing relevant temporal segments from video based on language query, has always been a highly challenging task in the field of video understanding. Videos often have a larger volume of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Feng Yue , Zhaoxing Zhang , Junming Jiao , Zhengyu Liang , Shiwen Cao , Feifei Zhang , Rong Shen

Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Xiaolong Sun , Le Wang , Sanping Zhou , Liushuai Shi , Kun Xia , Mengnan Liu , Yabing Wang , Gang Hua

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Ziheng Zhou , Jinxing Zhou , Wei Qian , Shengeng Tang , Xiaojun Chang , Dan Guo

We propose Context-aware Video-text Alignment (CVA), a novel framework to address a significant challenge in video temporal grounding: achieving temporally sensitive video-text alignment that remains robust to irrelevant background context.…

Machine Learning · Computer Science 2026-03-27 Sungho Moon , Seunghun Lee , Jiwan Seo , Sunghoon Im

Temporal action proposal generation (TAPG) aims to estimate temporal intervals of actions in untrimmed videos, which is a challenging yet plays an important role in many tasks of video analysis and understanding. Despite the great…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Khoa Vo , Kashu Yamazaki , Sang Truong , Minh-Triet Tran , Akihiro Sugimoto , Ngan Le

Understanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understanding Multi-modal Large Language Models (MLLMs) are capable of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Yingxian Chen , Jiahui Liu , Ruidi Fan , Yanwei Li , Chirui Chang , Shizhen Zhao , Wilton W. T. Fok , Xiaojuan Qi , Yik-Chung Wu

We introduce a novel deep learning-based audio-visual quality (AVQ) prediction model that leverages internal features from state-of-the-art unimodal predictors. Unlike prior approaches that rely on simple fusion strategies, our model…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Ina Salaj , Arijit Biswas

Audio-visual learning has been a major pillar of multi-modal machine learning, where the community mostly focused on its modality-aligned setting, i.e., the audio and visual modality are both assumed to signal the prediction target. With…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Yung-Hsuan Lai , Yen-Chun Chen , Yu-Chiang Frank Wang

We present CAT-V (Caption AnyThing in Video), a training-free framework for fine-grained object-centric video captioning that enables detailed descriptions of user-selected objects through time. CAT-V integrates three key components: a…

Many vision-language tasks can be reduced to the problem of sequence prediction for natural language output. In particular, recent advances in image captioning use deep reinforcement learning (RL) to alleviate the "exposure bias" during…

Computer Vision and Pattern Recognition · Computer Science 2018-08-23 Daqing Liu , Zheng-Jun Zha , Hanwang Zhang , Yongdong Zhang , Feng Wu

Temporal Video Grounding (TVG) aims to localize the temporal boundary of a specific segment in an untrimmed video based on a given language query. Since datasets in this domain are often gathered from limited video scenes, models tend to…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Haifeng Huang , Yang Zhao , Zehan Wang , Yan Xia , Zhou Zhao

Grounding temporal video segments described in natural language queries effectively and efficiently is a crucial capability needed in vision-and-language fields. In this paper, we deal with the fast video temporal grounding (FVTG) task,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-13 Ziyue Wu , Junyu Gao , Shucheng Huang , Changsheng Xu

We present AVID, the first large-scale benchmark for audio-visual inconsistency understanding in videos. While omni-modal large language models excel at temporally aligned tasks such as captioning and question answering, they struggle to…

Multimedia · Computer Science 2026-04-16 Zixuan Chen , Depeng Wang , Hao Lin , Li Luo , Ke Xu , Ya Guo , Huijia Zhu , Tanfeng Sun , Xinghao Jiang

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and more challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jinxing Zhou , Ziheng Zhou , Yanghao Zhou , Yuxin Mao , Zhangling Duan , Dan Guo

Temporal Action Localization (TAL) has garnered significant attention in information retrieval. Existing supervised or weakly supervised methods heavily rely on labeled temporal boundaries and action categories, which are labor-intensive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Rui Xia , Dan Jiang , Quan Zhang , Ke Zhang , Chun Yuan

Video moment retrieval (VMR) is to search for a visual temporal moment in an untrimmed raw video by a given text query description (sentence). Existing studies either start from collecting exhaustive frame-wise annotations on the temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Weitong Cai , Jiabo Huang , Shaogang Gong

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-15 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Hoirin Kim , Se-Young Yun
‹ Prev 1 3 4 5 6 7 10 Next ›