English
Related papers

Related papers: Contrastive Conditional Latent Diffusion for Audio…

200 papers

Existing audio-text retrieval (ATR) methods are essentially discriminative models that aim to maximize the conditional likelihood, represented as p(candidates|query). Nevertheless, this methodology fails to consider the intrinsic data…

Sound · Computer Science 2024-10-18 Yifei Xin , Xuxin Cheng , Zhihong Zhu , Xusheng Yang , Yuexian Zou

Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in webscale video datasets. However, conventional contrastive audio-visual…

Sound · Computer Science 2025-03-18 Ioannis Tsiamas , Santiago Pascual , Chunghsin Yeh , Joan Serrà

Learning robust representations for physiological time-series signals continues to pose a substantial challenge in developing efficient few-shot learning applications. This difficulty is largely due to the complex pathological variations in…

Machine Learning · Computer Science 2025-12-01 Rami Zewail

Community researchers have developed a range of advanced audio-visual segmentation models aimed at improving the quality of sounding objects' masks. While masks created by these models may initially appear plausible, they occasionally…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Peiwen Sun , Honggang Zhang , Di Hu

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or…

Sound · Computer Science 2025-01-20 Shengkui Zhao , Zexu Pan , Kun Zhou , Yukun Ma , Chong Zhang , Bin Ma

Recently, video recognition is emerging with the help of multi-modal learning, which focuses on integrating distinct modalities to improve the performance or robustness of the model. Although various multi-modal learning methods have been…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Haochen Han , Qinghua Zheng , Minnan Luo , Kaiyao Miao , Feng Tian , Yan Chen

Audiovisual segmentation (AVS) aims to identify visual regions corresponding to sound sources, playing a vital role in video understanding, surveillance, and human-computer interaction. Traditional AVS methods depend on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Seung-jae Lee , Paul Hongsuck Seo

The recent success of audio-visual representation learning can be largely attributed to their pervasive property of audio-visual synchronization, which can be used as self-annotated supervision. As a state-of-the-art solution, Audio-Visual…

Multimedia · Computer Science 2022-04-27 Hanyu Xuan , Yihong Xu , Shuo Chen , Zhiliang Wu , Jian Yang , Yan Yan , Xavier Alameda-Pineda

Self-supervised learning (SSL) approaches, such as contrastive and generative methods, have advanced environmental sound representation learning using unlabeled data. However, how these approaches can complement each other within a unified…

Sound · Computer Science 2025-10-29 Sivan Ding , Julia Wilkins , Magdalena Fuentes , Juan Pablo Bello

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data available; most audio-visual datasets are collected in…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-05 Ju-Chieh Chou , Chung-Ming Chien , Karen Livescu

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Sungnyun Kim , Sungwoo Cho , Sangmin Bae , Kangwook Jang , Se-Young Yun

Video saliency prediction is crucial for downstream applications, such as video compression and human-computer interaction. With the flourishing of multimodal learning, researchers started to explore multimodal video saliency prediction,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Li Yu , Xuanzhe Sun , Wei Zhou , Moncef Gabbouj

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform…

Recently, automatic speaker verification (ASV) based on deep learning is easily contaminated by adversarial attacks, which is a new type of attack that injects imperceptible perturbations to audio signals so as to make ASV produce wrong…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-10 Yibo Bai , Xiao-Lei Zhang , Xuelong Li

While conditional diffusion models are known to have good coverage of the data distribution, they still face limitations in output diversity, particularly when sampled with a high classifier-free guidance scale for optimal image quality or…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Seyedmorteza Sadat , Jakob Buhmann , Derek Bradley , Otmar Hilliges , Romann M. Weber

Blood vessel segmentation in medical imaging is one of the essential steps for vascular disease diagnosis and interventional planning in a broad spectrum of clinical scenarios in image-based medicine and interventional medicine.…

Image and Video Processing · Electrical Eng. & Systems 2023-08-02 Boah Kim , Yujin Oh , Bradford J. Wood , Ronald M. Summers , Jong Chul Ye

Existing machine learning research has achieved promising results in monaural audio-visual separation (MAVS). However, most MAVS methods purely consider what the sound source is, not where it is located. This can be a problem in VR/AR…

Sound · Computer Science 2023-11-01 Yuxin Ye , Wenming Yang , Yapeng Tian

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embeddings obtained from…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Julius Richter , Simone Frintrop , Timo Gerkmann

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual…

Computer Vision and Pattern Recognition · Computer Science 2020-11-04 Pedro Morgado , Yi Li , Nuno Vasconcelos