English
Related papers

Related papers: Automated Audio Captioning and Language-Based Audi…

200 papers

Automated audio captioning (AAC) is a novel task, where a method takes as an input an audio sample and outputs a textual description (i.e. a caption) of its contents. Most AAC methods are adapted from from image captioning of machine…

Sound · Computer Science 2020-10-22 An Tran , Konstantinos Drossos , Tuomas Virtanen

To address Task 5 in the Detection and Classification of Acoustic Scenes and Events (DCASE) 2018 challenge, in this paper, we propose an ensemble learning system. The proposed system consists of three different models, based on…

Audio and Speech Processing · Electrical Eng. & Systems 2018-12-13 Jeremy Chew , Yingxiang Sun , Lahiru Jayasinghe , Chau Yuen

Automatically creating the description of an image using any natural languages sentence like English is a very challenging task. It requires expertise of both image processing as well as natural language processing. This paper discuss about…

Computer Vision and Pattern Recognition · Computer Science 2018-10-03 Parth Shah , Vishvajit Bakarola , Supriya Pati

Recent progress in self-supervised or unsupervised machine learning has opened the possibility of building a full speech processing system from raw audio without using any textual representations or expert labels such as phonemes,…

Computation and Language · Computer Science 2022-10-31 Ewan Dunbar , Nicolas Hamilakis , Emmanuel Dupoux

Image retrieval using spoken language cues has emerged as a promising direction in multimodal perception, yet leveraging speech in multi-speaker scenarios remains challenging. We propose a novel Target Speaker Speech-Image Retrieval task…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-12 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Xinyue Song , Xianghu Yue

The ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-01 Ross Cutler , Ando Saabas , Tanel Parnamaa , Marju Purin , Hannes Gamper , Sebastian Braun , Karsten Sørensen , Robert Aichner

Automatic Audio Captioning (AAC) is the task that aims to describe an audio signal using natural language. AAC systems take as input an audio signal and output a free-form text sentence, called a caption. Evaluating such systems is not…

Sound · Computer Science 2022-11-17 Etienne Labbé , Thomas Pellegrini , Julien Pinquier

To better support information retrieval tasks such as web search and open-domain question answering, growing effort is made to develop retrieval-oriented language models, e.g., RetroMAE and many others. Most of the existing works focus on…

Computation and Language · Computer Science 2023-05-05 Shitao Xiao , Zheng Liu , Yingxia Shao , Zhao Cao

Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leaving the spatial information in multi-channel recordings…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-16 Chenxin Yu , Hao Ma , Xu Li , Xiao-Lei Zhang , Mingjie Shao , Chi Zhang , Xuelong Li

Image captioning and cross-modal retrieval are examples of tasks that involve the joint analysis of visual and linguistic information. In connection to remote sensing imagery, these tasks can help non-expert users in extracting relevant…

Computer Vision and Pattern Recognition · Computer Science 2024-02-12 João Daniel Silva , João Magalhães , Devis Tuia , Bruno Martins

Singing voices contain much richer information than common voices, including varied vocal and acoustic properties. However, current open-source audio-text datasets for singing voices capture only a narrow range of attributes and lack…

Computation and Language · Computer Science 2025-08-19 Hyunjong Ok , Jaeho Lee

Acoustic scene classification (ASC) is one of the most popular problems in the field of machine listening. The objective of this problem is to classify an audio clip into one of the predefined scenes using only the audio data. This problem…

This paper presents the details of Task 1: Acoustic Scene Classification in the DCASE 2020 Challenge. The task consists of two subtasks: classification of data from multiple devices, requiring good generalization properties, and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-03 Toni Heittola , Annamaria Mesaros , Tuomas Virtanen

Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text…

Sound · Computer Science 2024-03-06 Xuenan Xu , Zhiling Zhang , Zelin Zhou , Pingyue Zhang , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Language-queried audio source separation (LASS) is a new paradigm for computational auditory scene analysis (CASA). LASS aims to separate a target sound from an audio mixture given a natural language query, which provides a natural and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Xubo Liu , Qiuqiang Kong , Yan Zhao , Haohe Liu , Yi Yuan , Yuzhuo Liu , Rui Xia , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

End-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging acoustic information lost in the intermediate textual…

Achieving pronunciation proficiency in a second language (L2) remains a challenge, despite the development of Computer-Assisted Pronunciation Training (CAPT) systems. Traditional CAPT systems often provide unintuitive feedback that lacks…

Sound · Computer Science 2026-01-22 Hongfu Liu , Zhouying Cui , Xiangming Gu , Ye Wang

The aim of this project is to implement and design arobust synthetic speech classifier for the IEEE Signal ProcessingCup 2022 challenge. Here, we learn a synthetic speech attributionmodel using the speech generated from various…

This paper proposes a benchmark of submissions to Detection and Classification Acoustic Scene and Events 2021 Challenge (DCASE) Task 4 representing a sampling of the state-of-the-art in Sound Event Detection task. The submissions are…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-02 Francesca Ronchini , Romain Serizel

We propose EnCLAP, a novel framework for automated audio captioning. EnCLAP employs two acoustic representation models, EnCodec and CLAP, along with a pretrained language model, BART. We also introduce a new training objective called masked…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-01 Jaeyeon Kim , Jaeyoon Jung , Jinjoo Lee , Sang Hoon Woo