中文
相关论文

相关论文: EnCLAP: Combining Neural Audio Codec and Audio-Tex…

200 篇论文

There are a thousand ways to caption an image. Contrastive Language Pretraining (CLIP) on the other hand, works by mapping an image and its caption to a single vector -- limiting how well CLIP-like models can represent the diverse ways to…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Samuel Lavoie , Polina Kirichenko , Mark Ibrahim , Mahmoud Assran , Andrew Gordon Wilson , Aaron Courville , Nicolas Ballas

Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures…

音频与语音处理 · 电气工程与系统科学 2025-11-20 Mingyue Huo , Wei-Cheng Tseng , Yiwen Shao , Hao Zhang , Dong Yu

Neural speech models build deeply entangled internal representations, which capture a variety of features (e.g., fundamental frequency, loudness, syntactic category, or semantic content of a word) in a distributed encoding. This complexity…

计算与语言 · 计算机科学 2024-10-07 Hosein Mohebbi , Grzegorz Chrupała , Willem Zuidema , Afra Alishahi , Ivan Titov

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

声音 · 计算机科学 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

We introduce BANC, a neural binaural audio codec designed for efficient speech compression in single and two-speaker scenarios while preserving the spatial location information of each speaker. Our key contributions are as follows: 1) The…

声音 · 计算机科学 2024-11-26 Anton Ratnarajah , Shi-Xiong Zhang , Dong Yu

Automated Audio Captioning (AAC) generates captions for audio clips but faces challenges due to limited datasets compared to image captioning. To overcome this, we propose the zero-shot AAC system that leverages pre-trained models,…

计算与语言 · 计算机科学 2025-09-17 Vijay Govindarajan , Pratik Patel , Sahil Tripathi , Md Azizul Hoque , Gautam Siddharth Kashyap

Transformers rely on both content-based and position-based addressing mechanisms to make predictions, but existing positional encoding techniques often diminish the effectiveness of position-based addressing. Many current methods enforce…

计算与语言 · 计算机科学 2025-08-22 Jiajun Zhu , Peihao Wang , Ruisi Cai , Jason D. Lee , Pan Li , Zhangyang Wang

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are…

音频与语音处理 · 电气工程与系统科学 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Yukai Li , Mingjie Shao , Qiuqiang Kong , Ju Liu

Audio-Visual Video Parsing (AVVP) task aims to detect and temporally locate events within audio and visual modalities. Multiple events can overlap in the timeline, making identification challenging. While traditional methods usually focus…

人工智能 · 计算机科学 2024-07-12 Jinxing Zhou , Dan Guo , Yuxin Mao , Yiran Zhong , Xiaojun Chang , Meng Wang

This study introduces a novel training paradigm, audio difference learning, for improving audio captioning. The fundamental concept of the proposed learning method is to create a feature representation space that preserves the relationship…

音频与语音处理 · 电气工程与系统科学 2023-09-18 Tatsuya Komatsu , Yusuke Fujita , Kazuya Takeda , Tomoki Toda

Contrastive Language-Audio Pretraining (CLAP) models have demonstrated unprecedented performance in various acoustic signal recognition tasks. Fiber-optic-based acoustic recognition is one of the most important downstream tasks and plays a…

音频与语音处理 · 电气工程与系统科学 2025-01-20 Jingchen Sun , Shaobo Han , Wataru Kohno , Changyou Chen

Despite significant advancements in medical vision-language pre-training, existing methods have largely overlooked the inherent linguistic complexity and imbalanced isssue within medical reports, as well as the complex cross-modality…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Rongsheng Wang , Qingsong Yao , Zihang Jiang , Haoran Lai , Zhiyang He , Xiaodong Tao , S. Kevin Zhou

The goal of audio captioning is to translate input audio into its description using natural language. One of the problems in audio captioning is the lack of training data due to the difficulty in collecting audio-caption pairs by crawling…

音频与语音处理 · 电气工程与系统科学 2020-12-15 Yuma Koizumi , Yasunori Ohishi , Daisuke Niizumi , Daiki Takeuchi , Masahiro Yasuda

Automatic Audio Captioning (AAC) refers to the task of translating audio into a natural language that describes the audio events, source of the events and their relationships. The limited samples in AAC datasets at present, has set up a…

声音 · 计算机科学 2022-02-01 Swapnil Bhosale , Rupayan Chakraborty , Sunil Kumar Kopparapu

In this paper, we propose an algorithm, Epochal Difficult Captions, to supplement the training of any model for the Automated Audio Captioning task. Epochal Difficult Captions is an elegant evolution to the keyword estimation task that…

计算与语言 · 计算机科学 2022-06-07 Andrew Koh , Soham Tiwari , Chng Eng Siong

Large-scale vision-language models demonstrate strong multimodal alignment and generalization across diverse tasks. Among them, CLIP stands out as one of the most successful approaches. In this work, we extend the application of CLIP to…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Sooyoung Park , Arda Senocak , Joon Son Chung

Recent work has shown that generation from a prompted or fine-tuned language model can perform well at semantic parsing when the output is constrained to be a valid semantic representation. We introduce BenchCLAMP, a Benchmark to evaluate…

计算与语言 · 计算机科学 2024-01-11 Subhro Roy , Sam Thomson , Tongfei Chen , Richard Shin , Adam Pauls , Jason Eisner , Benjamin Van Durme

This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering)…

计算机视觉与模式识别 · 计算机科学 2019-12-05 Luowei Zhou , Hamid Palangi , Lei Zhang , Houdong Hu , Jason J. Corso , Jianfeng Gao

It is an open challenge to obtain high quality training data, especially captions, for text-to-audio models. Although prior methods have leveraged \textit{text-only language models} to augment and improve captions, such methods have…

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP's pretraining on…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Zhe Li , Weihao Yuan , Yisheng He , Lingteng Qiu , Shenhao Zhu , Xiaodong Gu , Weichao Shen , Yuan Dong , Zilong Dong , Laurence T. Yang