English
Related papers

Related papers: L3DAS22 Challenge: Learning 3D Audio Sources in a …

200 papers

During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not…

Sound · Computer Science 2022-07-22 Wei Sun , Mei Wang , Lili Qiu

This work describes our group's submission to the PROCESS Challenge 2024, with the goal of assessing cognitive decline through spontaneous speech, using three guided clinical tasks. This joint effort followed a holistic approach,…

In this paper, we present the task description and discuss the results of the DCASE 2020 Challenge Task 2: Unsupervised Detection of Anomalous Sounds for Machine Condition Monitoring. The goal of anomalous sound detection (ASD) is to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Yuma Koizumi , Yohei Kawaguchi , Keisuke Imoto , Toshiki Nakamura , Yuki Nikaido , Ryo Tanabe , Harsh Purohit , Kaori Suefusa , Takashi Endo , Masahiro Yasuda , Noboru Harada

This paper introduces the fifth oriental language recognition (OLR) challenge AP20-OLR, which intends to improve the performance of language recognition systems, along with APSIPA Annual Summit and Conference (APSIPA ASC). The data profile,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-12 Zheng Li , Miao Zhao , Qingyang Hong , Lin Li , Zhiyuan Tang , Dong Wang , Liming Song , Cheng Yang

Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab…

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Hashim Ali , Surya Subramani , Lekha Bollinani , Nithin Sai Adupa , Sali El-Loh , Hafiz Malik

This paper presents our MSXF TTS system for Task 3.1 of the Audio Deep Synthesis Detection (ADD) Challenge 2022. We use an end to end text to speech system, and add a constraint loss to the system when training stage. The end to end TTS…

Sound · Computer Science 2022-01-28 Chunyong Yang , Pengfei Liu , Yanli Chen , Hongbin Wang , Min Liu

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation…

Artificial Intelligence · Computer Science 2025-01-16 Mathieu Lagrange , Junwon Lee , Modan Tailleur , Laurie M. Heller , Keunwoo Choi , Brian McFee , Keisuke Imoto , Yuki Okamoto

This paper summarizes the outcomes from the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC). We first address the necessity of the challenge and then introduce the associated dataset collected from a new-energy vehicle…

Sound · Computer Science 2022-11-04 Ao Zhang , Fan Yu , Kaixun Huang , Lei Xie , Longbiao Wang , Eng Siong Chng , Hui Bu , Binbin Zhang , Wei Chen , Xin Xu

This report provides an overview of the challenge hosted at the OpenSUN3D Workshop on Open-Vocabulary 3D Scene Understanding held in conjunction with ICCV 2023. The goal of this workshop series is to provide a platform for exploration and…

Recent progress in audio generation models has made it possible to create highly realistic and immersive soundscapes, which are now widely used in film and virtual-reality-related applications. However, these audio generators also raise…

Sound · Computer Science 2026-01-01 Han Yin , Yang Xiao , Rohan Kumar Das , Jisheng Bai , Ting Dang

The AutoSpeech challenge calls for automated machine learning (AutoML) solutions to automate the process of applying machine learning to speech processing tasks. These tasks, which cover a large variety of domains, will be shown to the…

Artificial Intelligence · Computer Science 2020-10-27 Jingsong Wang , Tom Ko , Zhen Xu , Xiawei Guo , Souxiang Liu , Wei-Wei Tu , Lei Xie

The first Natural Office Talkers in Settings of Far-field Audio Recordings (NOTSOFAR-1) Challenge is a pivotal initiative that sets new benchmarks by offering datasets more representative of the needs of real-world business applications…

Sound · Computer Science 2025-03-11 Igor Abramovski , Alon Vinnikov , Shalev Shaer , Naoyuki Kanda , Xiaofei Wang , Amir Ivry , Eyal Krupka

This paper presents an analysis of the Low-Complexity Acoustic Scene Classification task in DCASE 2022 Challenge. The task was a continuation from the previous years, but the low-complexity requirements were changed to the following: the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-14 Irene Martín-Morató , Francesco Paissan , Alberto Ancilotto , Toni Heittola , Annamaria Mesaros , Elisabetta Farella , Alessio Brutti , Tuomas Virtanen

This paper describes LeVoice automatic speech recognition systems to track2 of intelligent cockpit speech recognition challenge 2022. Track2 is a speech recognition task without limits on the scope of model size. Our main points include…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-18 Yan Jia , Mi Hong , Jingyu Hou , Kailong Ren , Sifan Ma , Jin Wang , Fangzhen Peng , Yinglin Ji , Lin Yang , Junjie Wang

Recent advancements in large audio-language models (LALMs) have shown impressive capabilities in understanding and reasoning about audio and speech information. However, these models still face challenges, including hallucinating…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Chun-Yi Kuan , Hung-yi Lee

Data augmentation methods have shown great importance in diverse supervised learning problems where labeled data is scarce or costly to obtain. For sound event localization and detection (SELD) tasks several augmentation methods have been…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-20 Ricardo Falcon-Perez , Kazuki Shimada , Yuichiro Koyama , Shusuke Takahashi , Yuki Mitsufuji

The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition,…

The addition of Foley sound effects during post-production is a common technique used to enhance the perceived acoustic properties of multimedia content. Traditionally, Foley sound has been produced by human Foley artists, which involves…

This technical report details our systems submitted for Task 3 of the DCASE 2024 Challenge: Audio and Audiovisual Sound Event Localization and Detection (SELD) with Source Distance Estimation (SDE). We address only the audio-only SELD with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Jun Wei Yeow , Ee-Leng Tan , Jisheng Bai , Santi Peksi , Woon-Seng Gan
‹ Prev 1 3 4 5 6 7 10 Next ›