中文
相关论文

相关论文: Asteroid: the PyTorch-based audio source separatio…

200 篇论文

Monoaural audio source separation is a challenging research area in machine learning. In this area, a mixture containing multiple audio sources is given, and a model is expected to disentangle the mixture into isolated atomic sources. In…

机器学习 · 计算机科学 2019-11-25 Amir Zadeh , Tianjun Ma , Soujanya Poria , Louis-Philippe Morency

Pyserini is an easy-to-use Python toolkit that supports replicable IR research by providing effective first-stage retrieval in a multi-stage ranking architecture. Our toolkit is self-contained as a standard Python package and comes with…

信息检索 · 计算机科学 2021-02-22 Jimmy Lin , Xueguang Ma , Sheng-Chieh Lin , Jheng-Hong Yang , Ronak Pradeep , Rodrigo Nogueira

Microphone array post-filters have demonstrated their ability to greatly reduce noise at the output of a beamformer. However, current techniques only consider a single source of interest, most of the time assuming stationary background…

声音 · 计算机科学 2016-03-11 Jean-Marc Valin , Jean Rouat , François Michaud

Asteroseismology is well-established in astronomy as the gold standard for determining precise and accurate fundamental stellar properties like masses, radii, and ages. Several tools have been developed for asteroseismic analyses but many…

太阳与恒星天体物理 · 物理学 2021-09-08 Ashley Chontos , Daniel Huber , Maryum Sayeed , Pavadol Yamsiri

We propose a new dataset for cinematic audio source separation (CASS) that handles non-verbal sounds. Existing CASS datasets only contain reading-style sounds as a speech stem. These datasets differ from actual movie audio, which is more…

声音 · 计算机科学 2025-06-10 Takuya Hasumi , Yusuke Fujita

DORAEMON is an open-source PyTorch library that unifies visual object modeling and representation learning across diverse scales. A single YAML-driven workflow covers classification, retrieval and metric learning; more than 1000 pretrained…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Ke Du , Yimin Peng , Chao Gao , Fan Zhou , Siqiao Xue

This work presents a speech-to-text system "Pisets" for scientists and journalists which is based on a three-component architecture aimed at improving speech recognition accuracy while minimizing errors and hallucinations associated with…

计算与语言 · 计算机科学 2026-01-27 Ivan Bondarenko , Daniil Grebenkin , Oleg Sedukhin , Mikhail Klementev , Roman Derunets , Lyudmila Budneva

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Jiaben Chen , Renrui Zhang , Dongze Lian , Jiaqi Yang , Ziyao Zeng , Jianbo Shi

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and,…

音频与语音处理 · 电气工程与系统科学 2019-02-22 Thilo von Neumann , Keisuke Kinoshita , Marc Delcroix , Shoko Araki , Tomohiro Nakatani , Reinhold Haeb-Umbach

We study the single-channel source separation problem involving orthogonal frequency-division multiplexing (OFDM) signals, which are ubiquitous in many modern-day digital communication systems. Related efforts have been pursued in monaural…

信号处理 · 电气工程与系统科学 2023-06-28 Gary C. F. Lee , Amir Weiss , Alejandro Lancho , Yury Polyanskiy , Gregory W. Wornell

We explore active audio-visual separation for dynamic sound sources, where an embodied agent moves intelligently in a 3D environment to continuously isolate the time-varying audio stream being emitted by an object of interest. The agent…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Sagnik Majumder , Kristen Grauman

Recent advancements in Neural Audio Synthesis (NAS) have outpaced the development of standardized evaluation methodologies and tools. To bridge this gap, we introduce AquaTk, an open-source Python library specifically designed to simplify…

声音 · 计算机科学 2023-11-20 Ashvala Vinay , Alexander Lerch

The acoustic startle reflex (ASR) that can be induced by a loud sound stimulus can be used as a versatile tool to, e.g., estimate hearing thresholds or identify subjective tinnitus percepts in rodents. These techniques are based on the fact…

Recent years have seen immense progress in 3D computer vision and computer graphics, with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However, alongside immersive visual…

声音 · 计算机科学 2024-06-12 Mason Wang , Ryosuke Sawata , Samuel Clarke , Ruohan Gao , Shangzhe Wu , Jiajun Wu

N-HANS is a Python toolkit for in-the-wild audio enhancement, including speech, music, and general audio denoising, separation, and selective noise or source suppression. The functionalities are realised based on two neural network models…

声音 · 计算机科学 2019-12-02 Shuo Liu , Gil Keren , Björn Schuller

Most music source separation systems require large collections of isolated sources for training, which can be difficult to obtain. In this work, we use musical scores, which are comparatively easy to obtain, as a weak label for training a…

声音 · 计算机科学 2020-10-23 Yun-Ning Hung , Gordon Wichern , Jonathan Le Roux

With the continuous development of deep learning-based speech conversion and speech synthesis technologies, the cybersecurity problem posed by fake audio has become increasingly serious. Previously proposed models for defending against fake…

声音 · 计算机科学 2025-06-04 Chi Ding , Junxiao Xue , Cong Wang , Hao Zhou

We introduce PyChEst, a Python package which provides tools for the simultaneous estimation of multiple changepoints in the distribution of piece-wise stationary time series. The nonparametric algorithms implemented are provably consistent…

统计计算 · 统计学 2021-12-21 Azadeh Khaleghi , Lukas Zierahn

Microgrids, self contained electrical grids that are capable of disconnecting from the main grid, hold potential in both tackling climate change mitigation via reducing CO2 emissions and adaptation by increasing infrastructure resiliency.…

人工智能 · 计算机科学 2020-11-17 Gonzague Henri , Tanguy Levent , Avishai Halev , Reda Alami , Philippe Cordier

Music source separation is the task of isolating the instrumental tracks from a music song. Despite its spectacular recent progress, the trend towards more complex architectures and training protocols exacerbates reproducibility issues. The…

声音 · 计算机科学 2026-03-11 Paul Magron , Romain Serizel , Constance Douwes