中文
相关论文

相关论文: TONet: Tone-Octave Network for Singing Melody Extr…

200 篇论文

With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in order to explore how…

声音 · 计算机科学 2021-08-29 Dengfeng Ke , Yuxing Lu , Xudong Liu , Yanyan Xu , Jing Sun , Cheng-Hao Cai

This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2…

声音 · 计算机科学 2025-04-15 Yubing Cao , Yinfeng Yu , Yongming Li , Liejun Wang

In artificial-intelligence-aided signal processing, existing deep learning models often exhibit a black-box structure, and their validity and comprehensibility remain elusive. The integration of topological methods, despite its relatively…

机器学习 · 计算机科学 2023-11-28 Pingyao Feng , Siheng Yi , Qingrui Qu , Zhiwang Yu , Yifei Zhu

In this paper, we propose a novel approach for the optimal identification of correlated segments in noisy correlation matrices. The proposed model is known as CoSeNet (Correlation Seg-mentation Network) and is based on a four-layer…

Identifying musical instruments in polyphonic music recordings is a challenging but important problem in the field of music information retrieval. It enables music search by instrument, helps recognize musical genres, or can make music…

声音 · 计算机科学 2016-12-28 Yoonchang Han , Jaehun Kim , Kyogu Lee

Music source separation aims to separate polyphonic music into different types of sources. Most existing methods focus on enhancing the quality of separated results by using a larger model structure, rendering them unsuitable for deployment…

声音 · 计算机科学 2024-07-02 Chun-Hsiang Wang , Chung-Che Wang , Jun-You Wang , Jyh-Shing Roger Jang , Yen-Hsun Chu

Speaker extraction is to extract a target speaker's voice from multi-talker speech. It simulates humans' cocktail party effect or the selective listening ability. The prior work mostly performs speaker extraction in frequency domain, then…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio…

音频与语音处理 · 电气工程与系统科学 2024-02-13 Kavya Ranjan Saxena , Vipul Arora

The challenges of polyphonic sound event detection (PSED) stem from the detection of multiple overlapping events in a time series. Recent efforts exploit Deep Neural Networks (DNNs) on Time-Frequency Representations (TFRs) of audio clips as…

声音 · 计算机科学 2021-11-29 Wangkai Jin , Junyu Liu , Jianfeng Ren , Xiangjun Peng

Optical Music Recognition (OMR) is an important technology within Music Information Retrieval. Deep learning models show promising results on OMR tasks, but symbol-level annotated data sets of sufficient size to train such models are not…

计算机视觉与模式识别 · 计算机科学 2017-07-18 Eelco van der Wel , Karen Ullrich

Singing voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we…

声音 · 计算机科学 2022-05-25 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Occlusion edge detection requires both accurate locations and context constraints of the contour. Existing CNN-based pipeline does not utilize adaptive methods to filter the noise introduced by low-level features. To address this dilemma,…

计算机视觉与模式识别 · 计算机科学 2019-03-22 Rui Lu , Menghan Zhou , Anlong Ming , Yu Zhou

Melody preservation is crucial in singing voice conversion (SVC). However, in many scenarios, audio is often accompanied with background music (BGM), which can cause audio distortion and interfere with the extraction of melody and other key…

声音 · 计算机科学 2025-02-10 Wei Chen , Binzhu Sha , Jing Yang , Zhuo Wang , Fan Fan , Zhiyong Wu

A new musical instrument classification method using convolutional neural networks (CNNs) is presented in this paper. Unlike the traditional methods, we investigated a scheme for classifying musical instruments using the learned features…

声音 · 计算机科学 2015-12-24 Taejin Park , Taejin Lee

Onsets are a key factor to split audio into several notes. In this paper, we ensemble multiple temporal convolution network (TCN) based model and utilize a restricted frequency range spectrogram to achieve more robust onset detection.…

声音 · 计算机科学 2023-06-09 Yu Cheng Hung , Jian-Jiun Ding

Lyrics transcription of polyphonic music is challenging because singing vocals are corrupted by the background music. To improve the robustness of lyrics transcription to the background music, we propose a strategy of combining the features…

音频与语音处理 · 电气工程与系统科学 2022-04-25 Xiaoxue Gao , Chitralekha Gupta , Haizhou Li

Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform…

声音 · 计算机科学 2024-05-31 Adam Sorrenti

Efficient and accurate bird sound classification is of important for ecology, habitat protection and scientific research, as it plays a central role in monitoring the distribution and abundance of species. However, prevailing methods…

声音 · 计算机科学 2023-12-27 Yiyuan Yang , Kaichen Zhou , Niki Trigoni , Andrew Markham

Recent studies show that self-attentions behave like low-pass filters (as opposed to convolutions) and enhancing their high-pass filtering capability improves model performance. Contrary to this idea, we investigate existing…

计算机视觉与模式识别 · 计算机科学 2023-08-23 Guhnoo Yun , Juhan Yoo , Kijung Kim , Jeongho Lee , Dong Hwan Kim

Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods…

声音 · 计算机科学 2025-10-13 Ke Xue , Rongfei Fan , Lixin , Dawei Zhao , Chao Zhu , Han Hu