中文
相关论文

相关论文: AID: Open-source Anechoic Interferer Dataset

200 篇论文

The imitation of percussive instruments via the human voice is a natural way for us to communicate rhythmic ideas and, for this reason, it attracts the interest of music makers. Specifically, the automatic mapping of these vocal imitations…

音频与语音处理 · 电气工程与系统科学 2020-09-25 Alejandro Delgado , SKoT McDonald , Ning Xu , Mark Sandler

This paper describes an open-source Python framework for handling datasets for music processing tasks, built with the aim of improving the reproducibility of research projects in music computing and assessing the generalization abilities of…

多媒体 · 计算机科学 2021-12-28 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution…

计算与语言 · 计算机科学 2024-08-13 Ruibo Liu , Jerry Wei , Fangyu Liu , Chenglei Si , Yanzhe Zhang , Jinmeng Rao , Steven Zheng , Daiyi Peng , Diyi Yang , Denny Zhou , Andrew M. Dai

Privacy-preserving crowd density analysis finds application across a wide range of scenarios, substantially enhancing smart building operation and management while upholding privacy expectations in various spaces. We propose a non-speech…

We present pyroomacoustics, a software package aimed at the rapid development and testing of audio array processing algorithms. The content of the package can be divided into three main components: an intuitive Python object-oriented…

声音 · 计算机科学 2019-05-08 Robin Scheibler , Eric Bezzam , Ivan Dokmanić

Traditional disaster analysis and modelling tools for assessing the severity of a disaster are predictive in nature. Based on the past observational data, these tools prescribe how the current input state (e.g., environmental conditions,…

Open data refers to data that is freely available for reuse. Although there has been rapid increase in availability of open data to public in the last decade, this has not translated into better decision-support tools for them. We propose…

人工智能 · 计算机科学 2019-01-14 Biplav Srivastava

Neurodivergent people frequently experience decreased sound tolerance, with estimates suggesting it affects 50-70% of this population. This heightened sensitivity can provoke reactions ranging from mild discomfort to severe distress,…

音频与语音处理 · 电气工程与系统科学 2025-09-15 Alexander Popescu , Rosie Frost , Milos Cernak

Artificial intelligence (AI) is transforming supply chain management, yet progress in predictive tasks -- such as delivery delay prediction -- remains constrained by the scarcity of high-quality, openly available datasets. Existing datasets…

人工智能 · 计算机科学 2025-09-09 Liming Xu , Yunbo Long , Alexandra Brintrup

Pythonic code is idiomatic code that follows guiding principles and practices within the Python community. Offering performance and readability benefits, Pythonic code is claimed to be widely adopted by experienced Python developers, but…

We present Binaspect, an open-source Python library for binaural audio analysis, visualization, and feature generation. Binaspect generates interpretable "azimuth maps" by calculating modified interaural time and level difference…

声音 · 计算机科学 2025-10-30 Dan Barry , Davoud Shariat Panah , Alessandro Ragano , Jan Skoglund , Andrew Hines

In this note we consider the problem of synthesizing optimal control policies for a system from noisy datasets. We present a novel algorithm that takes as input the available dataset and, based on these inputs, computes an optimal policy…

最优化与控制 · 数学 2020-03-02 Davide Gagliardi , Giovanni Russo

We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It consists of subsystems for signal synchronization,…

音频与语音处理 · 电气工程与系统科学 2022-05-03 Tobias Gburrek , Christoph Boeddeker , Thilo von Neumann , Tobias Cord-Landwehr , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

We present an analytic provenance data repository that can be used to study human analysis activity, thought processes, and software interaction with visual analysis tools during exploratory data analysis. We conducted a series of user…

人机交互 · 计算机科学 2018-01-17 Sina Mohseni , Andrew Pachuilo , Ehsanul Haque Nirjhar , Rhema Linder , Alyssa Pena , Eric D. Ragan

This paper presents BUT ReverbDB - a dataset of real room impulse responses (RIR), background noises and re-transmitted speech data. The retransmitted data includes LibriSpeech test-clean, 2000 HUB5 English evaluation and part of 2010 NIST…

音频与语音处理 · 电气工程与系统科学 2019-09-04 Igor Szoke , Miroslav Skacel , Ladislav Mosner , Jakub Paliesek , Jan "Honza" Cernocky

Heart and lung sounds are crucial for healthcare monitoring. Recent improvements in stethoscope technology have made it possible to capture patient sounds with enhanced precision. In this dataset, we used a digital stethoscope to capture…

音频与语音处理 · 电气工程与系统科学 2026-05-15 Yasaman Torabi , Shahram Shirani , James P. Reilly

SpeechPy is an open source Python package that contains speech preprocessing techniques, speech features, and important post-processing operations. It provides most frequent used speech features including MFCCs and filterbank energies…

声音 · 计算机科学 2018-07-25 Amirsina Torfi

The paper introduces a novel methodology for the identification of coefficients of switched autoregressive linear models. We consider the case when the system's outputs are contaminated by possibly large values of measurement noise. It is…

系统与控制 · 计算机科学 2019-03-27 Sarah Hojjatinia , Constantino M. Lagoa , Fabrizio Dabbene

Advancements in artificial intelligence and machine learning have significantly improved synthetic speech generation. This paper explores diffusion models, a novel method for creating realistic synthetic speech. We create a diffusion…

密码学与安全 · 计算机科学 2025-01-15 Anton Firc , Kamil Malinka , Petr Hanáček

The effectiveness of existing denoising algorithms typically relies on accurate pre-defined noise statistics or plenty of paired data, which limits their practicality. In this work, we focus on denoising in the more common case where noise…

图像与视频处理 · 电气工程与系统科学 2020-12-01 Huangxing Lin , Yihong Zhuang , Yue Huang , Xinghao Ding , Yizhou Yu , Xiaoqing Liu , John Paisley