中文
相关论文

相关论文: Bootstrapping deep music separation from primitive…

200 篇论文

With the recent advancements of data driven approaches using deep neural networks, music source separation has been formulated as an instrument-specific supervised problem. While existing deep learning models implicitly absorb the spatial…

音频与语音处理 · 电气工程与系统科学 2022-02-16 Darius Petermann , Minje Kim

Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extracted high-level semantics from visual inputs as the guidance…

声音 · 计算机科学 2024-07-08 Shentong Mo , Yapeng Tian

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

This work focuses on reliable detection of bird sound emissions as recorded in the open field. Acoustic detection of avian sounds can be used for the automatized monitoring of multiple bird taxa and querying in long-term recordings for…

声音 · 计算机科学 2016-09-28 Ilyas Potamitis

With increasing amounts of music being digitally transferred from production to distribution, automatic means of determining media quality are needed. Protection mechanisms in digital audio processing tools have not eliminated the need of…

声音 · 计算机科学 2022-02-14 Daniel Wolff , Rémi Mignot , Axel Roebel

The task of manipulating the level and/or effects of individual instruments to recompose a mixture of recordings, or remixing, is common across a variety of applications such as music production, audio-visual post-production, podcasts, and…

音频与语音处理 · 电气工程与系统科学 2021-10-25 Haici Yang , Shivani Firodiya , Nicholas J. Bryan , Minje Kim

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti

Producing manual, pixel-accurate, image segmentation labels is tedious and time-consuming. This is often a rate-limiting factor when large amounts of labeled images are required, such as for training deep convolutional networks for…

计算机视觉与模式识别 · 计算机科学 2021-02-19 Luis C. Garcia-Peraza-Herrera , Lucas Fidon , Claudia D'Ettorre , Danail Stoyanov , Tom Vercauteren , Sebastien Ourselin

In music source separation, a standard training data augmentation procedure is to create new training samples by randomly combining instrument stems from different songs. These random mixes have mismatched characteristics compared to real…

音频与语音处理 · 电气工程与系统科学 2024-02-29 Chang-Bin Jeon , Gordon Wichern , François G. Germain , Jonathan Le Roux

Music source separation has been a popular topic in signal processing for decades, not only because of its technical difficulty, but also due to its importance to many commercial applications, such as automatic karoake and remixing. In this…

音频与语音处理 · 电气工程与系统科学 2020-03-23 Yuzhou Liu , Balaji Thoshkahna , Ali Milani , Trausti Kristjansson

Separating audio mixtures into individual instrument tracks has been a long standing challenging task. We introduce a novel weakly supervised audio source separation approach based on deep adversarial learning. Specifically, our loss…

声音 · 计算机科学 2018-05-18 Ning Zhang , Junchi Yan , Yuchen Zhou

Speech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail party scenarios. As a…

声音 · 计算机科学 2024-11-06 Wupeng Wang , Zexu Pan , Xinke Li , Shuai Wang , Haizhou Li

Segmenting audio into homogeneous sections such as music and speech helps us understand the content of audio. It is useful as a pre-processing step to index, store, and modify audio recordings, radio broadcasts and TV programmes. Deep…

Choral singing is a widely practiced form of ensemble singing wherein a group of people sing simultaneously in polyphonic harmony. The most commonly practiced setting for choir ensembles consists of four parts; Soprano, Alto, Tenor and Bass…

音频与语音处理 · 电气工程与系统科学 2020-08-19 Darius Petermann , Pritish Chandna , Helena Cuesta , Jordi Bonada , Emilia Gomez

Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text…

声音 · 计算机科学 2024-03-06 Xuenan Xu , Zhiling Zhang , Zelin Zhou , Pingyue Zhang , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Existing automatic music generation approaches that feature deep learning can be broadly classified into two types: raw audio models and symbolic models. Symbolic models, which train and generate at the note level, are currently the more…

声音 · 计算机科学 2018-06-27 Rachel Manzelli , Vijay Thakkar , Ali Siahkamari , Brian Kulis

The proliferation of distorted, compressed, and manipulated music on modern media platforms like TikTok motivates the development of more robust audio fingerprinting techniques to identify the sources of musical recordings. In this paper,…

声音 · 计算机科学 2025-11-10 Shubhr Singh , Kiran Bhat , Xavier Riley , Benjamin Resnick , John Thickstun , Walter De Brouwer

The audio source separation tasks, such as speech enhancement, speech separation, and music source separation, have achieved impressive performance in recent studies. The powerful modeling capabilities of deep neural networks give us hope…

音频与语音处理 · 电气工程与系统科学 2021-07-15 Lu Zhang , Chenxing Li , Feng Deng , Xiaorui Wang

Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack…

音频与语音处理 · 电气工程与系统科学 2025-02-18 Jaime Garcia-Martinez , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen , Julio J. Carabias-Orti , Pedro Vera-Candeas

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representation using…

计算机视觉与模式识别 · 计算机科学 2016-10-31 Yusuf Aytar , Carl Vondrick , Antonio Torralba