English
Related papers

Related papers: Audio Captioning with Composition of Acoustic and …

200 papers

With the advent of modern AI architectures, a shift has happened towards end-to-end architectures. This pivot has led to neural architectures being trained without domain-specific biases/knowledge, optimized according to the task. We in…

Sound · Computer Science 2025-05-08 Prateek Verma

This paper proposes a zero-shot learning approach for audio classification based on the textual information about class labels without any audio samples from target classes. We propose an audio classification system built on the bilinear…

Machine Learning · Computer Science 2019-08-08 Huang Xie , Tuomas Virtanen

Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio, different people may perceive the same audio differently, resulting in caption disparities (i.e., one audio may correlate to…

Sound · Computer Science 2022-04-19 Yiming Zhang , Hong Yu , Ruoyi Du , Zhanyu Ma , Yuan Dong

This work presents a text-to-audio-retrieval system based on pre-trained text and spectrogram transformers. Our method projects recordings and textual descriptions into a shared audio-caption space in which related examples from different…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-09 Paul Primus , Khaled Koutini , Gerhard Widmer

Automated audio captioning aims at generating textual descriptions for an audio clip. To evaluate the quality of generated audio captions, previous works directly adopt image captioning metrics like SPICE and CIDEr, without justifying their…

Sound · Computer Science 2022-01-28 Zelin Zhou , Zhiling Zhang , Xuenan Xu , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when…

Sound · Computer Science 2025-01-07 Xiquan Li , Wenxi Chen , Ziyang Ma , Xuenan Xu , Yuzhe Liang , Zhisheng Zheng , Qiuqiang Kong , Xie Chen

While existing speech audio codecs designed for compression exploit limited forms of temporal redundancy and allow for multi-scale representations, they tend to represent all features of audio in the same way. In contrast, generative voice…

Sound · Computer Science 2025-09-22 Ryan Collette , Ross Greenwood , Serena Nicoll

The Automated Audio Captioning (AAC) task aims to describe an audio signal using natural language. To evaluate machine-generated captions, the metrics should take into account audio events, acoustic scenes, paralinguistics, signal…

Sound · Computer Science 2024-11-06 Satvik Dixit , Soham Deshmukh , Bhiksha Raj

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy…

Graphics · Computer Science 2021-12-02 Seung Hyun Lee , Wonseok Roh , Wonmin Byeon , Sang Ho Yoon , Chan Young Kim , Jinkyu Kim , Sangpil Kim

Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This…

Computation and Language · Computer Science 2025-01-07 Ariel Shaulov , Tal Shaharabany , Eitan Shaar , Gal Chechik , Lior Wolf

Automated audio captioning (AAC) is the task of automatically generating textual descriptions for general audio signals. A captioning system has to identify various information from the input signal and express it with natural language.…

Machine Learning · Computer Science 2021-10-15 Benno Weck , Xavier Favory , Konstantinos Drossos , Xavier Serra

This paper proposes to use similarities of audio captions for estimating audio-caption relevances to be used for training text-based audio retrieval systems. Current audio-caption datasets (e.g., Clotho) contain audio samples paired with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Huang Xie , Khazar Khorrami , Okko Räsänen , Tuomas Virtanen

The explosion of video data on the internet requires effective and efficient technology to generate captions automatically for people who are not able to watch the videos. Despite the great progress of video captioning research,…

Computer Vision and Pattern Recognition · Computer Science 2018-07-11 Xiangxi Shi , Jianfei Cai , Jiuxiang Gu , Shafiq Joty

Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-27 Soumya Dutta , Sriram Ganapathy

Automated audio captioning (AAC), a task that mimics human perception as well as innovatively links audio processing and natural language processing, has overseen much progress over the last few years. AAC requires recognizing contents such…

Sound · Computer Science 2023-11-17 Xuenan Xu , Zeyu Xie , Mengyue Wu , Kai Yu

Automated audio captioning (AAC) aims to describe the content of an audio clip using simple sentences. Existing AAC methods are developed based on an encoder-decoder architecture that success is attributed to the use of a pre-trained CNN10…

Sound · Computer Science 2022-10-18 Jianyuan Sun , Xubo Liu , Xinhao Mei , Mark D. Plumbley , Volkan Kilic , Wenwu Wang

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining…

Sound · Computer Science 2024-10-11 Alican Akman , Qiyang Sun , Björn W. Schuller

In recent years, advancements in representation learning and language models have propelled Automated Captioning (AC) to new heights, enabling the generation of human-level descriptions. Leveraging these advancements, we propose AVCap, an…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-12 Jongsuk Kim , Jiwon Shin , Junmo Kim

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding…

Sound · Computer Science 2021-10-12 Cheng Gong , Longbiao Wang , Zhenhua Ling , Ju Zhang , Jianwu Dang

Accurately detecting emotions in conversation is a necessary yet challenging task due to the complexity of emotions and dynamics in dialogues. The emotional state of a speaker can be influenced by many different factors, such as…

Computation and Language · Computer Science 2023-02-07 Jiachen Luo , Huy Phan , Joshua Reiss