中文
相关论文

相关论文: An Explainable Proxy Model for Multiabel Audio Seg…

200 篇论文

Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, as they do not utilize existing knowledge of how sound is…

机器学习 · 计算机科学 2020-01-15 Jesse Engel , Lamtharn Hantrakul , Chenjie Gu , Adam Roberts

Deep neural networks (DNNs) have achieved substantial predictive performance in various speech processing tasks. Particularly, it has been shown that a monaural speech separation task can be successfully solved with a DNN-based method…

音频与语音处理 · 电气工程与系统科学 2021-04-20 Chihiro Watanabe , Hirokazu Kameoka

Label noise is a critical problem in medical image segmentation, often arising from the inherent difficulty of manual annotation. Models trained on noisy data are prone to overfitting, which degrades their generalization performance. While…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Wesam Moustafa , Hossam Elsafty , Helen Schneider , Lorenz Sparrenberg , Rafet Sifa

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

音频与语音处理 · 电气工程与系统科学 2025-01-16 Amit Eliav , Sharon Gannot

High quality labeled datasets have allowed deep learning to achieve impressive results on many sound analysis tasks. Yet, it is labor-intensive to accurately annotate large amount of audio data, and the dataset may contain noisy labels in…

音频与语音处理 · 电气工程与系统科学 2020-07-17 Boqing Zhu , Kele Xu , Qiuqiang Kong , Huaimin Wang , Yuxing Peng

Explaining the decisions made by audio spoofing detection models is crucial for fostering trust in detection outcomes. However, current research on the interpretability of detection models is limited to applying XAI tools to post-trained…

声音 · 计算机科学 2025-07-28 Menglu Li , Xiao-Ping Zhang

A recent line of research on automated speaking assessment (ASA) has benefited from self-supervised learning (SSL) representations, which capture rich acoustic and linguistic patterns in non-native speech without underlying assumptions of…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Tien-Hong Lo , Szu-Yu Chen , Yao-Ting Sung , Berlin Chen

Unsupervised anomalous sound detection (ASD) aims to identify anomalous sounds by learning the features of normal operational sounds and sensing their deviations. Recent approaches have focused on the self-supervised task utilizing the…

声音 · 计算机科学 2023-10-11 Soonhyeon Choi , Jung-Woo Choi

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a two-stage pipeline -…

声音 · 计算机科学 2025-07-24 Tobias Morocutti , Jonathan Greif , Paul Primus , Florian Schmid , Gerhard Widmer

Semantic segmentation is essential in computer vision for various applications, yet traditional approaches face significant challenges, including the high cost of annotation and extensive training for supervised learning. Additionally, due…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Yasufumi Kawano , Yoshimitsu Aoki

After constructing a deep neural network for urban sound classification, this work focuses on the sensitive application of assisting drivers suffering from hearing loss. As such, clear etiology justifying and interpreting model predictions…

声音 · 计算机科学 2021-11-22 Marco Colussi , Stavros Ntalampiras

Denoising diffusion probabilistic models have recently received much research attention since they outperform alternative approaches, such as GANs, and currently provide state-of-the-art generative performance. The superior performance of…

计算机视觉与模式识别 · 计算机科学 2022-03-17 Dmitry Baranchuk , Ivan Rubachev , Andrey Voynov , Valentin Khrulkov , Artem Babenko

Machine anomalous sound detection (ASD) is a valuable technique across various applications. However, its generalization performance is often limited due to challenges in data collection and the complexity of acoustic environments. Inspired…

声音 · 计算机科学 2025-08-19 Bing Han , Anbai Jiang , Xinhu Zheng , Wei-Qiang Zhang , Jia Liu , Pingyi Fan , Yanmin Qian

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker…

声音 · 计算机科学 2024-12-25 Xuechen Liu , Junichi Yamagishi , Md Sahidullah , Tomi kinnunen

In recent years Artificial Intelligence has emerged as a fundamental tool in medical applications. Despite this rapid development, deep neural networks remain black boxes that are difficult to explain, and this represents a major limitation…

图像与视频处理 · 电气工程与系统科学 2024-05-22 Tommaso Torda , Andrea Ciardiello , Simona Gargiulo , Greta Grillo , Simone Scardapane , Cecilia Voena , Stefano Giagu

As deep vision models' popularity rapidly increases, there is a growing emphasis on explanations for model predictions. The inherently explainable attribution method aims to enhance the understanding of model behavior by identifying the…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Xianren Zhang , Dongwon Lee , Suhang Wang

Recently, neural networks based purely on self-attention, such as the Vision Transformer (ViT), have been shown to outperform deep learning models constructed with convolutional neural networks (CNNs) on various vision tasks, thus extending…

声音 · 计算机科学 2022-02-14 Yuan Gong , Cheng-I Jeff Lai , Yu-An Chung , James Glass

Building generalizable AI models is one of the primary challenges in the healthcare domain. While radiologists rely on generalizable descriptive rules of abnormality, Neural Network (NN) models suffer even with a slight shift in input…

计算机视觉与模式识别 · 计算机科学 2023-07-12 Shantanu Ghosh , Ke Yu , Kayhan Batmanghelich

Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first introduce speaker…

声音 · 计算机科学 2023-12-19 Peng Shen , Xugang Lu , Hisashi Kawai

This paper proposes a novel Sequence-to-Sequence Neural Diarization (S2SND) framework to perform online and offline speaker diarization. It is developed from the sequence-to-sequence architecture of our previous target-speaker voice…

音频与语音处理 · 电气工程与系统科学 2025-06-24 Ming Cheng , Yuke Lin , Ming Li