English
Related papers

Related papers: CAPTDURE: Captioned Sound Dataset of Single Source…

200 papers

We live in a rich and varied acoustic world, which is experienced by individuals or communities as a soundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-26 Yuanbo Hou , Qiaoqiao Ren , Andrew Mitchell , Wenwu Wang , Jian Kang , Tony Belpaeme , Dick Botteldooren

Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This…

Computation and Language · Computer Science 2025-01-07 Ariel Shaulov , Tal Shaharabany , Eitan Shaar , Gal Chechik , Lior Wolf

Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-03 Paul Primus , Gerhard Widmer

Speech separation is very important in real-world applications such as human-machine interaction, hearing aids devices, and automatic meeting transcription. In recent years, a significant improvement occurred towards the solution based on…

Sound · Computer Science 2024-08-29 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text --…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Karan Desai , Gaurav Kaul , Zubin Aysola , Justin Johnson

Environmental sound synthesis is a technique for generating a natural environmental sound. Conventional work on environmental sound synthesis using sound event labels cannot finely control synthesized sounds, for example, the pitch and…

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional…

Computation and Language · Computer Science 2018-04-11 Pete Warden

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

Training large vision-language models requires extensive, high-quality image-text pairs. Existing web-scraped datasets, however, are noisy and lack detailed image descriptions. To bridge this gap, we introduce PixelProse, a comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Vasu Singla , Kaiyu Yue , Sukriti Paul , Reza Shirkavand , Mayuka Jayawardhana , Alireza Ganjdanesh , Heng Huang , Abhinav Bhatele , Gowthami Somepalli , Tom Goldstein

Cinematic audio source separation (CASS), as a problem of extracting the dialogue, music, and effects stems from their mixture, is a relatively new subtask of audio source separation. To date, only one publicly available dataset exists for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 Karn N. Watcharasupat , Chih-Wei Wu , Iroro Orife

Humans use context to assess the veracity of information. However, current audio deepfake detectors only analyze the audio file without considering either context or transcripts. We create and analyze a Journalist-provided Deepfake Dataset…

Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS approaches are audio-only, overlooking the inherent…

Multimedia · Computer Science 2026-03-30 Kang Zhang , Suyeon Lee , Arda Senocak , Joon Son Chung

Automated audio captioning is a cross-modal translation task for describing the content of audio clips with natural language sentences. This task has attracted increasing attention and substantial progress has been made in recent years.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-02 Xinhao Mei , Xubo Liu , Jianyuan Sun , Mark D. Plumbley , Wenwu Wang

Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely…

The performance of single channel source separation algorithms has improved greatly in recent times with the development and deployment of neural networks. However, many such networks continue to operate on the magnitude spectrogram of a…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-08 Shrikant Venkataramani , Paris Smaragdis

Target Sound Extraction (TSE) focuses on the problem of separating sources of interest, indicated by a user's cue, from the input mixture. Most existing solutions operate in an offline fashion and are not suited to the low-latency causal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-22 Shrishail Baligar , Mikolaj Kegler , Bryce Irvin , Marko Stamenovic , Shawn Newsam

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial information (e.g.,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Tuochao Chen , D Shin , Hakan Erdogan , Sinan Hersek

Fully-supervised models for source separation are trained on parallel mixture-source data and are currently state-of-the-art. However, such parallel data is often difficult to obtain, and it is cumbersome to adapt trained models to mixtures…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-30 Ge Zhu , Jordan Darefsky , Fei Jiang , Anton Selitskiy , Zhiyao Duan

Conventional approaches to sound localization and separation are based on microphone arrays in artificial systems. Inspired by the selective perception of human auditory system, we design a multi-source listening system which can separate…

Sound · Computer Science 2019-11-11 Xuecong Sun , Han Jia , Zhe Zhang , Yuzhen Yang , Zhaoyong Sun , Jun Yang

This work presents a text-to-audio-retrieval system based on pre-trained text and spectrogram transformers. Our method projects recordings and textual descriptions into a shared audio-caption space in which related examples from different…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-09 Paul Primus , Khaled Koutini , Gerhard Widmer