English
Related papers

Related papers: Hyperbolic Audio Source Separation

200 papers

Localized point sources (monopoles) in an acoustical domain are implemented to a three dimensional non-singular Helmholtz boundary element method in the frequency domain. It allows for the straightforward use of higher order surface…

Numerical Analysis · Mathematics 2023-05-04 Qiang Sun

We propose a hyperbolic set-to-set distance measure for computing dissimilarity between sets in hyperbolic space. While point-to-point distances in hyperbolic space effectively capture hierarchical relationships between data points, many…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Pengxiang Li , Wei Wu , Zhi Gao , Xiaomeng Fan , Peilin Yu , Yuwei Wu , Zhipeng Lu , Yunde Jia , Mehrtash Harandi

Retrieval-augmented generation (RAG) for biomedical knowledge faces a hierarchy-aware ontology grounding challenge: resources like HPO, DO, and MeSH use deep ``is-a" taxonomies, yet production stacks rely on Euclidean embeddings and ANN…

Information Retrieval · Computer Science 2026-04-14 Ou Deng , Shoji Nishimura , Atsushi Ogihara , Qun Jin

Hyperspectral unmixing is a blind source separation problem which consists in estimating the reference spectral signatures contained in a hyperspectral image, as well as their relative contribution to each pixel according to a given mixture…

Data Analysis, Statistics and Probability · Physics 2017-11-21 Pierre-Antoine Thouvenin , Nicolas Dobigeon , Jean-Yves Tourneret

In this work, we propose a hierarchical subspace model for acoustic unit discovery. In this approach, we frame the task as one of learning embeddings on a low-dimensional phonetic subspace, and simultaneously specify the subspace itself as…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-10 Bolaji Yusuf , Lucas Ondel , Lukas Burget , Jan Cernocky , Murat Saraclar

In this thesis, we propose an artificial auditory system that gives a robot the ability to locate and track sounds, as well as to separate simultaneous sound sources and recognising simultaneous speech. We demonstrate that it is possible to…

Robotics · Computer Science 2016-02-23 Jean-Marc Valin

Recently, hyperspherical embeddings have established themselves as a dominant technique for face and voice recognition. Specifically, Euclidean space vector embeddings are learned to encode person-specific information in their direction…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-25 Nikita Kuzmin , Igor Fedorov , Alexey Sholokhov

Recently, significant progress has been made in audio source separation by the application of deep learning techniques. Current methods that combine both audio and visual information use 2D representations such as images to guide the…

Sound · Computer Science 2021-02-04 Francesc Lluís , Vasileios Chatziioannou , Alex Hofmann

We propose an explainable probabilistic framework for characterizing spoofed speech by decomposing it into probabilistic attribute embeddings. Unlike raw high-dimensional countermeasure embeddings, which lack interpretability, the proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Jagabandhu Mishra , Manasi Chhibber , Hye-jin Shim , Tomi H. Kinnunen

We propose a visually conditioned music remixing system by incorporating deep visual and audio models. The method is based on a state of the art audio-visual source separation model which performs music instrument source separation with…

Sound · Computer Science 2020-10-29 Li-Chia Yang , Alexander Lerch

Incremental improvements in accuracy of Convolutional Neural Networks are usually achieved through use of deeper and more complex models trained on larger datasets. However, enlarging dataset and models increases the computation and storage…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-24 Mahdi Hajibabaei , Dengxin Dai

We are looking for a mathematical model of monophonic sounds with independent time and phase dimensions. With such a model we can resynthesise a sound with arbitrarily modulated frequency and progress of the timbre. We propose such a model…

Sound · Computer Science 2011-03-22 Henning Thielemann

In neural-based audio feature extraction, ensuring that representations capture disentangled information is crucial for model interpretability. However, existing disentanglement methods often rely on assumptions that are highly dependent on…

Sound · Computer Science 2025-10-07 Benoit Ginies , Xiaoyu Bie , Olivier Fercoq , Gaël Richard

Upsampling artifacts are caused by problematic upsampling layers and due to spectral replicas that emerge while upsampling. Also, depending on the used upsampling layer, such artifacts can either be tonal artifacts (additive high-frequency…

Sound · Computer Science 2021-11-24 Jordi Pons , Joan Serrà , Santiago Pascual , Giulio Cengarle , Daniel Arteaga , Davide Scaini

In this work, we demonstrate how a publicly available, pre-trained Jukebox model can be adapted for the problem of audio source separation from a single mixed audio channel. Our neural network architecture, which is using transfer learning,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-22 W. Zai El Amri , O. Tautz , H. Ritter , A. Melnik

Automated bioacoustic analysis aids understanding and protection of both marine and terrestrial animals and their habitats across extensive spatiotemporal scales, and typically involves analyzing vast collections of acoustic data. With the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-22 Burooj Ghani , Tom Denton , Stefan Kahl , Holger Klinck

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Shota Horiguchi , Takafumi Moriya , Atsushi Ando , Takanori Ashihara , Hiroshi Sato , Naohiro Tawara , Marc Delcroix

This research paper presents a novel deep learning-based neural network architecture, named Y-Net, for achieving music source separation. The proposed architecture performs end-to-end hybrid source separation by extracting features from…

We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding…

Sound · Computer Science 2021-02-11 Zeqian Li , Jacob Whitehill

Electroencephalography (EEG)-based brain-computer interfaces facilitate direct communication with a computer, enabling promising applications in human-computer interactions. However, their utility is currently limited because EEG decoding…

Machine Learning · Computer Science 2026-02-10 Shanglin Li , Shiwen Chu , Okan Koç , Yi Ding , Qibin Zhao , Motoaki Kawanabe , Ziheng Chen