English
Related papers

Related papers: TEn-CATG:Text-Enriched Audio-Visual Video Parsing …

200 papers

Weakly supervised audio-visual video parsing (AVVP) methods aim to detect audible-only, visible-only, and audible-visible events using only video-level labels. Existing approaches tackle this by leveraging unimodal and cross-modal contexts.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Faegheh Sardari , Armin Mustafa , Philip J. B. Jackson , Adrian Hilton

Audio-Visual Question Answering (AVQA) task aims to answer questions about different visual objects, sounds, and their associations in videos. Such naturally multi-modal videos are composed of rich and complex dynamic audio-visual…

Computer Vision and Pattern Recognition · Computer Science 2023-08-11 Guangyao Li , Wenxuan Hou , Di Hu

Dense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data.…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Wei-Yuan Cheng , Kai-Po Chang , Chi-Pin Huang , Fu-En Yang , Yu-Chiang Frank Wang

We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Jinxing Zhou , Jianyuan Wang , Jiayi Zhang , Weixuan Sun , Jing Zhang , Stan Birchfield , Dan Guo , Lingpeng Kong , Meng Wang , Yiran Zhong

Mobile Edge Caching (MEC) is a revolutionary technology for the Sixth Generation (6G) of wireless networks with the promise to significantly reduce users' latency via offering storage capacities at the edge of the network. The efficiency of…

Machine Learning · Computer Science 2022-10-28 Zohreh HajiAkhondi-Meybodi , Arash Mohammadi , Ming Hou , Jamshid Abouei , Konstantinos N. Plataniotis

We propose a new task named Audio-driven Per-formance Video Generation (APVG), which aims to synthesizethe video of a person playing a certain instrument guided bya given music audio clip. It is a challenging task to gener-ate the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-06 Hao Zhu , Yi Li , Feixia Zhu , Aihua Zheng , Ran He

With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the complex relationship…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Wenjing Wang , Huan Yang , Zixi Tuo , Huiguo He , Junchen Zhu , Jianlong Fu , Jiaying Liu

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a challenge, in part…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Edson Araujo , Saurabhchand Bhati , M. Jehanzeb Mirza , Brian Kingsbury , Samuel Thomas , Rogerio Feris , James R. Glass , Hilde Kuehne

Audio-Visual Segmentation (AVS) aims to precisely outline audible objects in a visual scene at the pixel level. Existing AVS methods require fine-grained annotations of audio-mask pairs in supervised learning fashion. This limits their…

Computer Vision and Pattern Recognition · Computer Science 2023-09-14 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Xiatian Zhu

Detecting text in natural scenes remains challenging, particularly for diverse scripts and arbitrarily shaped instances where visual cues alone are often insufficient. Existing methods do not fully leverage semantic context. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Mohammed-En-Nadhir Zighem , Abdenour Hadid

Training deep learning models for video classification from audio-visual data commonly requires immense amounts of labeled training data collected via a costly process. A challenging and underexplored, yet much cheaper, setup is few-shot…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Otniel-Bogdan Mercea , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

Weakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Jie Fu , Junyu Gao , Changsheng Xu

Endeavors have been made to explore Large Language Models for video analysis (Video-LLMs), particularly in understanding and interpreting long videos. However, existing Video-LLMs still face challenges in effectively integrating the rich…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Jungang Li , Sicheng Tao , Yibo Yan , Xiaojie Gu , Haodong Xu , Xu Zheng , Yuanhuiyi Lyu , Linfeng Zhang , Xuming Hu

Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intricate local structures…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Shangzhe Di , Zhonghua Zhai , Weidi Xie

A promising approach for steering auditory attention in complex listening environments relies on Auditory Attention Decoding (AAD), which aim to identify the attended speech stream in a multiple speaker scenario from neural recordings.…

Clustering high-dimensional multivariate spatiotemporal climate data is challenging due to complex temporal dependencies, evolving spatial interactions, and non-stationary dynamics. Conventional clustering methods, including recurrent and…

Machine Learning · Computer Science 2025-09-17 Francis Ndikum Nji , Vandana Janaja , Jianwu Wang

Multi-label image classification is a prediction task that aims to identify more than one label from a given image. This paper considers the semantic consistency of the latent space between the visual patch and linguistic label domains and…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Miaoge Li , Dongsheng Wang , Xinyang Liu , Zequn Zeng , Ruiying Lu , Bo Chen , Mingyuan Zhou

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Moayed Haji-Ali , Willi Menapace , Aliaksandr Siarohin , Ivan Skorokhodov , Alper Canberk , Kwot Sin Lee , Vicente Ordonez , Sergey Tulyakov

Semantic segmentation is a fundamental task in computer vision that involves dense pixel-wise classification for scene understanding. Despite significant progress, achieving high accuracy while maintaining real-time performance remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Abhinav Sagar

We propose a weakly-supervised framework for action labeling in video, where only the order of occurring actions is required during training time. The key challenge is that the per-frame alignments between the input (video) and label…

Computer Vision and Pattern Recognition · Computer Science 2016-07-29 De-An Huang , Li Fei-Fei , Juan Carlos Niebles