English
Related papers

Related papers: A decade of DCASE: Achievements, practices, evalua…

200 papers

PLACES 2013 (full title: Programming Language Approaches to Concurrency- and Communication-cEntric Software) was the sixth edition of the PLACES workshop series. After the first PLACES, which was affiliated to DisCoTec in 2008, the workshop…

Programming Languages · Computer Science 2013-12-10 Nobuko Yoshida , Wim Vanderbauwhede

In this paper, we describe in detail our systems for DCASE 2020 Task 4. The systems are based on the 1st-place system of DCASE 2019 Task 4, which adopts weakly-supervised framework with an attention-based embedding-level pooling module and…

Sound · Computer Science 2020-11-03 Yuxin Huang , Liwei Lin , Shuo Ma , Xiangdong Wang , Hong Liu , Yueliang Qian , Min Liu , Kazushige Ouch

Decentralized learning offers a promising approach to crowdsource data consumptions and computational workloads across geographically distributed compute interconnected through peer-to-peer networks, accommodating the exponentially…

Machine Learning · Computer Science 2025-07-10 Tongtian Zhu , Wenhao Li , Can Wang , Fengxiang He

In this paper, we present a deep neural network (DNN)-based acoustic scene classification framework. Two hierarchical learning methods are proposed to improve the DNN baseline performance by incorporating the hierarchical taxonomy…

Sound · Computer Science 2016-08-16 Yong Xu , Qiang Huang , Wenwu Wang , Mark D. Plumbley

The "VOiCES from a Distance Challenge 2019" is designed to foster research in the area of speaker recognition and automatic speech recognition (ASR) with the special focus on single channel distant/far-field audio, under noisy conditions.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-01 Mahesh Kumar Nandwana , Julien van Hout , Mitchell McLaren , Colleen Richey , Aaron Lawson , Maria Alejandra Barrios

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-25 Hongyu Wang , Chenda Li , Xin Zhou , Shuai Wang , Yanmin Qian

Audio captioning is an important research area that aims to generate meaningful descriptions for audio clips. Most of the existing research extracts acoustic features of audio clips as input to encoder-decoder and transformer architectures…

Sound · Computer Science 2022-04-20 Ayşegül Özkaya Eren , Mustafa Sert

Recently, more and more personalized speech enhancement systems (PSE) with excellent performance have been proposed. However, two critical issues still limit the performance and generalization ability of the model: 1) Acoustic environment…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Xiaofeng Ge , Jiangyu Han , Haixin Guan , Yanhua Long

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP…

Polyphonic sound event detection and direction-of-arrival estimation require different input features from audio signals. While sound event detection mainly relies on time-frequency patterns, direction-of-arrival estimation relies on…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-17 Thi Ngoc Tho Nguyen , Douglas L. Jones , Woon-Seng Gan

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting…

Sound · Computer Science 2025-10-28 Pengfei Cai , Yan Song , Qing Gu , Nan Jiang , Haoyu Song , Ian McLoughlin

This paper summarizes the ICASSP 2026 Automatic Song Aesthetics Evaluation (ASAE) Challenge, which focuses on predicting the subjective aesthetic scores of AI-generated songs. The challenge consists of two tracks: Track 1 targets the…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Guobin Ma , Yuxuan Xia , Jixun Yao , Huixin Xue , Hexin Liu , Shuai Wang , Hao Liu , Lei Xie

This technical report describes the IOA team's submission for TASK1A of DCASE2019 challenge. Our acoustic scene classification (ASC) system adopts a data augmentation scheme employing generative adversary networks. Two major classifiers, 1D…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-17 Hangting Chen , Zuozhen Liu , Zongming Liu , Pengyuan Zhang , Yonghong Yan

Acoustic Scene Classification (ASC) identifies an environment based on an audio signal. This paper explores ASC in low-resource conditions and proposes a novel model, DS-FlexiNet, which combines depthwise separable convolutions from…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-29 Zhi Chen , Yun-Fei Shao , Yong Ma , Mingsheng Wei , Le Zhang , Wei-Qiang Zhang

Acoustic Scene Classification (ASC) is one of the core research problems in the field of Computational Sound Scene Analysis. In this work, we present SubSpectralNet, a novel model which captures discriminative features by incorporating…

Sound · Computer Science 2019-02-26 Sai Samarth R Phaye , Emmanouil Benetos , Ye Wang

This report presents the dataset and baseline of Task 3 of the DCASE2021 Challenge on Sound Event Localization and Detection (SELD). The dataset is based on emulation of real recordings of static or moving sound events under real conditions…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-06 Archontis Politis , Sharath Adavanne , Daniel Krause , Antoine Deleforge , Prerak Srivastava , Tuomas Virtanen

Acoustic scene classification is an intricate problem for a machine. As an emerging field of research, deep Convolutional Neural Networks (CNN) achieve convincing results. In this paper, we explore the use of multi-scale Dense connected…

Computer Vision and Pattern Recognition · Computer Science 2018-06-13 Dawei Feng , Kele Xu , Haibo Mi , Feifan Liao , Yan Zhou

Scene classification is a fundamental task in interpretation of remote sensing images, and has become an active research topic in remote sensing community due to its important role in a wide range of applications. Over the past years,…

Computer Vision and Pattern Recognition · Computer Science 2018-06-05 Fan Hu , Gui-Song Xia , Wen Yang , Liangpei Zhang

The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-16 Eric Guizzo , Christian Marinoni , Marco Pennese , Xinlei Ren , Xiguang Zheng , Chen Zhang , Bruno Masiero , Aurelio Uncini , Danilo Comminiello

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various…

Sound · Computer Science 2025-11-03 Jiarong Du , Zhan Jin , Peijun Yang , Juan Liu , Zhuo Li , Xin Liu , Ming Li
‹ Prev 1 8 9 10 Next ›