English
Related papers

Related papers: Description on IEEE ICME 2024 Grand Challenge: Sem…

200 papers

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module,…

Acoustic events are sounds with well-defined spectro-temporal characteristics which can be associated with the physical objects generating them. Acoustic scenes are collections of such acoustic events in no specific temporal order. Given…

Sound · Computer Science 2022-06-28 Rahil Parikh , Harshavardhan Sundar , Ming Sun , Chao Wang , Spyros Matsoukas

Labelled data are limited and self-supervised learning is one of the most important approaches for reducing labelling requirements. While it has been extensively explored in the image domain, it has so far not received the same amount of…

Sound · Computer Science 2025-05-16 Jingyong Liang , Bernd Meyer , Isaac Ning Lee , Thanh-Toan Do

We present the 2017 Visual Domain Adaptation (VisDA) dataset and challenge, a large-scale testbed for unsupervised domain adaptation across visual domains. Unsupervised domain adaptation aims to solve the real-world problem of domain shift,…

Computer Vision and Pattern Recognition · Computer Science 2017-11-30 Xingchao Peng , Ben Usman , Neela Kaushik , Judy Hoffman , Dequan Wang , Kate Saenko

We present an iVector based Acoustic Scene Classification (ASC) system suited for real life settings where active foreground speech can be present. In the proposed system, each recording is represented by a fixed-length iVector that models…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-03 Siyuan Song , Brecht Desplanques , Celest De Moor , Kris Demuynck , Nilesh Madhu

When performing data classification over a stream of continuously occurring instances, a key challenge is to develop an open-world classifier that anticipates instances from an unknown class. Studies addressing this problem, typically…

Computer Vision and Pattern Recognition · Computer Science 2018-10-10 Yang Gao , Swarup Chandra , Zhuoyi Wang , Latifur Khan

We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bioacoustics, Temporal Soundscapes, and Complex QA) to test…

The URGENT 2024 Challenge aims to foster speech enhancement (SE) techniques with great universality, robustness, and generalizability, featuring a broader task definition, large-scale multi-domain data, and comprehensive evaluation metrics.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Wangyou Zhang , Kohei Saijo , Samuele Cornell , Robin Scheibler , Chenda Li , Zhaoheng Ni , Anurag Kumar , Marvin Sach , Wei Wang , Yihui Fu , Shinji Watanabe , Tim Fingscheidt , Yanmin Qian

Self-supervised learning approaches have lately achieved great success on a broad spectrum of machine learning problems. In the field of speech processing, one of the most successful recent self-supervised models is wav2vec 2.0. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-10 Marie Kunešová , Zbyněk Zajíc

Recently, more and more personalized speech enhancement systems (PSE) with excellent performance have been proposed. However, two critical issues still limit the performance and generalization ability of the model: 1) Acoustic environment…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-23 Xiaofeng Ge , Jiangyu Han , Haixin Guan , Yanhua Long

Previous work has shown that it is possible to improve speech recognition by learning acoustic features from paired acoustic-articulatory data, for example by using canonical correlation analysis (CCA) or its deep extensions. One limitation…

Computation and Language · Computer Science 2018-03-21 Qingming Tang , Weiran Wang , Karen Livescu

Semi-supervised domain adaptation (SSDA) presents a critical hurdle in computer vision, especially given the frequent scarcity of labeled data in real-world settings. This scarcity often causes foundation models, trained on extensive…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Ali Mottaghi , Mohammad Abdullah Jamal , Serena Yeung , Omid Mohareri

This paper presents the details of the Audio-Visual Scene Classification task in the DCASE 2021 Challenge (Task 1 Subtask B). The task is concerned with classification using audio and video modalities, using a dataset of synchronized…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-21 Shanshan Wang , Toni Heittola , Annamaria Mesaros , Tuomas Virtanen

Sound event detection (SED) entails identifying the type of sound and estimating its temporal boundaries from acoustic signals. These events are uniquely characterized by their spatio-temporal features, which are determined by the way they…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Tanmay Khandelwal , Rohan Kumar Das

Cross-domain retrieval (CDR), as a crucial tool for numerous technologies, is finding increasingly broad applications. However, existing efforts face several major issues, with the most critical being the need for accurate supervision,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Lixu Wang , Xinyu Du , Qi Zhu

While there have been considerable advancements in machine learning driven by extensive datasets, a significant disparity still persists in the availability of data across various sources and populations. This inequality across domains…

Machine Learning · Computer Science 2024-03-11 Jinha Park , Wonguk Cho , Taesup Kim

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remove these background sounds using speech enhancement or train…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-04 Chaitanya Narisetty , Emiru Tsunoo , Xuankai Chang , Yosuke Kashiwagi , Michael Hentschel , Shinji Watanabe

During the last half decade, convolutional neural networks (CNNs) have triumphed over semantic segmentation, which is one of the core tasks in many applications such as autonomous driving. However, to train CNNs requires a considerable…

Computer Vision and Pattern Recognition · Computer Science 2018-11-15 Yang Zhang , Philip David , Boqing Gong

In recent years, the need for semantic segmentation has arisen across several different applications and environments. However, the expense and redundancy of annotation often limits the quantity of labels available for training in any…

Computer Vision and Pattern Recognition · Computer Science 2019-09-25 Tarun Kalluri , Girish Varma , Manmohan Chandraker , C V Jawahar

Environmental Sound Classification (ESC) is a challenging field of research in non-speech audio processing. Most of current research in ESC focuses on designing deep models with special architectures tailored for specific audio datasets,…

Sound · Computer Science 2021-03-03 Alireza Nasiri , Jianjun Hu
‹ Prev 1 8 9 10 Next ›