中文
相关论文

相关论文: CL4AC: A Contrastive Loss for Audio Captioning

200 篇论文

Neural audio codecs (NACs) achieve low-bitrate compression by learning compact audio representations, which can also serve as features for perceptual quality evaluation. We introduce DACe, an enhanced, higher-fidelity version of the…

音频与语音处理 · 电气工程与系统科学 2026-01-23 Arijit Biswas , Lars Villemoes

The dichotomy between the challenging nature of obtaining annotations for activities, and the more straightforward nature of data collection from wearables, has resulted in significant interest in the development of techniques that utilize…

机器学习 · 计算机科学 2022-11-14 Harish Haresamudram , Irfan Essa , Thomas Ploetz

Machine hearing of the environmental sound is one of the important issues in the audio recognition domain. It gives the machine the ability to discriminate between the different input sounds that guides its decision making. In this work we…

声音 · 计算机科学 2022-07-20 Peter Ochieng , Dennis Kaburu

Many speech enhancement methods try to learn the relationship between noisy and clean speech, obtained using an acoustic room simulator. We point out several limitations of enhancement methods relying on clean speech targets; the goal of…

计算与语言 · 计算机科学 2018-12-26 Geonmin Kim , Hwaran Lee , Bo-Kyeong Kim , Sang-Hoon Oh , Soo-Young Lee

Audio-visual captioning aims to generate holistic scene descriptions by jointly modeling sound and vision. While recent methods have improved performance through sophisticated modality fusion, it remains unclear to what extent the two…

音频与语音处理 · 电气工程与系统科学 2025-10-29 Yuchi Ishikawa , Toranosuke Manabe , Tatsuya Komatsu , Yoshimitsu Aoki

Audio Packet Loss Concealment (PLC) is the hiding of gaps in audio streams caused by data transmission failures in packet switched networks. This is a common problem, and of increasing importance as end-to-end VoIP telephony and…

声音 · 计算机科学 2022-04-12 Lorenz Diener , Sten Sootla , Solomiya Branets , Ando Saabas , Robert Aichner , Ross Cutler

In this paper, we first extend the recent Masked Auto-Encoder (MAE) model from a single modality to audio-visual multi-modalities. Subsequently, we propose the Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE) by combining contrastive…

Cross-modal retrieval (CMR) has been extensively applied in various domains, such as multimedia search engines and recommendation systems. Most existing CMR methods focus on image-to-text retrieval, whereas audio-to-text retrieval, a less…

声音 · 计算机科学 2023-09-19 Kaiyi Luo , Xulong Zhang , Jianzong Wang , Huaxiong Li , Ning Cheng , Jing Xiao

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely…

音频与语音处理 · 电气工程与系统科学 2025-05-01 Sungnyun Kim , Sungwoo Cho , Sangmin Bae , Kangwook Jang , Se-Young Yun

In recent years, user-generated audio content has proliferated across various media platforms, creating a growing need for efficient retrieval methods that allow users to search for audio clips using natural language queries. This task,…

声音 · 计算机科学 2024-12-31 Haoran Sun , Zimu Wang , Qiuyi Chen , Jianjun Chen , Jia Wang , Haiyang Zhang

The recent emergence of joint CTC-Attention model shows significant improvement in automatic speech recognition (ASR). The improvement largely lies in the modeling of linguistic information by decoder. The decoder joint-optimized with an…

计算与语言 · 计算机科学 2022-10-27 Xulong Zhang , Jianzong Wang , Ning Cheng , Mengyuan Zhao , Zhiyong Zhang , Jing Xiao

Acoustic scene classification (ASC) is a crucial research problem in computational auditory scene analysis, and it aims to recognize the unique acoustic characteristics of an environment. One of the challenges of the ASC task is the domain…

音频与语音处理 · 电气工程与系统科学 2024-03-01 Jisheng Bai , Mou Wang , Haohe Liu , Han Yin , Yafei Jia , Siwei Huang , Yutong Du , Dongzhe Zhang , Dongyuan Shi , Woon-Seng Gan , Mark D. Plumbley , Susanto Rahardja , Bin Xiang , Jianfeng Chen

Context compression is an advanced technique that accelerates large language model (LLM) inference by converting long inputs into compact representations. Existing methods primarily rely on autoencoding tasks to train special compression…

计算与语言 · 计算机科学 2026-03-12 Xin Liu , Runsong Zhao , Pengcheng Huang , Xinyu Liu , Junyi Xiao , Chunyang Xiao , Tong Xiao , Shengxiang Gao , Zhengtao Yu , Jingbo Zhu

For the Facial Action Unit (AU) detection task, accurately capturing the subtle facial differences between distinct AUs is essential for reliable detection. Additionally, AU detection faces challenges from class imbalance and the presence…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Ziqiao Shang , Bin Liu , Fengmao Lv , Fei Teng , Tianrui Li , Lan-Zhe Guo

Audio-text retrieval systems based on Contrastive Language-Audio Pretraining (CLAP) achieve strong performance on traditional benchmarks; however, these benchmarks rely on caption-style queries that differ substantially from real-world…

声音 · 计算机科学 2026-04-21 HaeJun Yoo , Yongseop Shin , Insung Lee , Myoung-Wan Koo , Du-Seong Chang

3D captioning, which aims to describe the content of 3D scenes in natural language, remains highly challenging due to the inherent sparsity of point clouds and weak cross-modal alignment in existing methods. To address these challenges, we…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Ting Huang , Zeyu Zhang , Yemin Wang , Hao Tang

Content-based music information retrieval has seen rapid progress with the adoption of deep learning. Current approaches to high-level music description typically make use of classification models, such as in auto-tagging or genre and mood…

声音 · 计算机科学 2021-12-09 Ilaria Manco , Emmanouil Benetos , Elio Quinton , Gyorgy Fazekas

This paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then…

Acoustic echo cancellation (AEC) aims to remove interference signals while leaving near-end speech least distorted. As the indistinguishable patterns between near-end speech and interference signals, near-end speech can't be separated…

音频与语音处理 · 电气工程与系统科学 2023-07-27 Chang Han , Xinmeng Xu , Weiping Tu , Yuhong Yang , Yajie Liu

The ability to efficiently search for images is essential for improving the user experiences across various products. Incorporating user feedback, via multi-modal inputs, to navigate visual search can help tailor retrieved results to…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Surgan Jandial , Pinkesh Badjatiya , Pranit Chawla , Ayush Chopra , Mausoom Sarkar , Balaji Krishnamurthy