中文
相关论文

相关论文: Deep Autotuner: a Pitch Correcting Network for Sin…

200 篇论文

Neural network based speech dereverberation has achieved promising results in recent studies. Nevertheless, many are focused on recovery of only the direct path sound and early reflections, which could be beneficial to speech perception,…

声音 · 计算机科学 2021-10-19 Ziteng Wang , Yueyue Na , Biao Tian , Qiang Fu

Study Objectives: Sleep stage scoring is performed manually by sleep experts and is prone to subjective interpretation of scoring rules with low intra- and interscorer reliability. Many automatic systems rely on few small-scale databases…

计算机视觉与模式识别 · 计算机科学 2020-08-24 Alexander Neergaard Olesen , Poul Jennum , Emmanuel Mignot , Helge B D Sorensen

Sound processing in the human auditory system is complex and highly non-linear, whereas hearing aids (HAs) still rely on simplified descriptions of auditory processing or hearing loss to restore hearing. Even though standard HA…

音频与语音处理 · 电气工程与系统科学 2023-06-21 Fotios Drakopoulos , Sarah Verhulst

Pitch is a foundational aspect of our perception of audio signals. Pitch contours are commonly used to analyze speech and music signals and as input features for many audio tasks, including music transcription, singing voice synthesis, and…

音频与语音处理 · 电气工程与系统科学 2024-08-13 Max Morrison , Caedon Hsieh , Nathan Pruyne , Bryan Pardo

Mood recognition is an important problem in music informatics and has key applications in music discovery and recommendation. These applications have become even more relevant with the rise of music streaming. Our work investigates the…

声音 · 计算机科学 2021-10-12 Rajnish Kumar , Manjeet Dahiya

Voice conversion is a task to convert a non-linguistic feature of a given utterance. Since naturalness of speech strongly depends on its pitch pattern, in some applications, it would be desirable to keep the original rise/fall pitch pattern…

音频与语音处理 · 电气工程与系统科学 2022-10-21 Chihiro Watanabe , Hirokazu Kameoka

There are many use cases in singing synthesis where creating voices from small amounts of data is desirable. In text-to-speech there have been several promising results that apply voice cloning techniques to modern deep learning based…

声音 · 计算机科学 2019-02-21 Merlijn Blaauw , Jordi Bonada , Ryunosuke Daido

Recursion is a fundamental concept in the design of filters and audio systems. In particular, artificial reverberation systems that use delay networks depend on recursive paths to control both echo density and the decay rate of modal…

音频与语音处理 · 电气工程与系统科学 2026-04-28 Gloria Dal Santo , Karolina Prawda , Sebastian J. Schlecht , Vesa Välimäki

Self-supervision methods learn representations by solving pretext tasks that do not require human-generated labels, alleviating the need for time-consuming annotations. These methods have been applied in computer vision, natural language…

声音 · 计算机科学 2023-06-27 Giovana Morais , Matthew E. P. Davies , Marcelo Queiroz , Magdalena Fuentes

Singing voice conversion is converting the timbre in the source singing to the target speaker's voice while keeping singing content the same. However, singing data for target speaker is much more difficult to collect compared with normal…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Liqiang Zhang , Chengzhu Yu , Heng Lu , Chao Weng , Chunlei Zhang , Yusong Wu , Xiang Xie , Zijin Li , Dong Yu

Pitch detection is a fundamental problem in speech processing as F0 is used in a large number of applications. Recent articles have proposed deep learning for robust pitch tracking. In this paper, we consider voicing detection as a…

声音 · 计算机科学 2019-03-06 Thomas Drugman , Goeric Huybrechts , Viacheslav Klimkov , Alexis Moinet

Finetuning (pretrained) language models is a standard approach for updating their internal parametric knowledge and specializing them to new tasks and domains. However, the corresponding model weight changes ("weight diffs") are not…

机器学习 · 计算机科学 2026-03-24 Avichal Goel , Yoon Kim , Nir Shavit , Tony T. Wang

The performance of robots in high-level tasks depends on the quality of their lower-level controller, which requires fine-tuning. However, the intrinsically nonlinear dynamics and controllers make tuning a challenging task when it is done…

机器人学 · 计算机科学 2024-07-12 Sheng Cheng , Minkyung Kim , Lin Song , Chengyu Yang , Yiquan Jin , Shenlong Wang , Naira Hovakimyan

Label noise in datasets could significantly damage the performance and robustness of deep neural networks (DNNs) trained on these datasets. As the size of modern DNNs grows, there is a growing demand for automated tools for detecting such…

机器学习 · 计算机科学 2025-10-28 Dang Huu-Tien , Minh-Phuong Nguyen , Naoya Inoue

Since the vocal component plays a crucial role in popular music, singing voice detection has been an active research topic in music information retrieval. Although several proposed algorithms have shown high performances, we argue that…

声音 · 计算机科学 2018-06-05 Kyungyun Lee , Keunwoo Choi , Juhan Nam

Despite the ubiquity of mobile and wearable text messaging applications, the problem of keyboard text decoding is not tackled sufficiently in the light of the enormous success of the deep learning Recurrent Neural Network (RNN) and…

计算与语言 · 计算机科学 2017-09-20 Shaona Ghosh , Per Ola Kristensson

Pitch estimation is an essential step of many speech processing algorithms, including speech coding, synthesis, and enhancement. Recently, pitch estimators based on deep neural networks (DNNs) have have been outperforming well-established…

音频与语音处理 · 电气工程与系统科学 2024-01-17 Krishna Subramani , Jean-Marc Valin , Jan Buethe , Paris Smaragdis , Mike Goodwin

The quantity of processed data is crucial for advancing the field of singing voice synthesis. While there are tools available for lyric or note transcription tasks, they all need pre-processed data which is relatively time-consuming (e.g.,…

声音 · 计算机科学 2024-10-11 Siwei Wu , Jinzheng He , Ruibin Yuan , Haojie Wei , Xipin Wei , Chenghua Lin , Jin Xu , Junyang Lin

Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone…

声音 · 计算机科学 2025-07-08 Nhan Duc Thanh Nguyen , Huy Phan , Simon Geirnaert , Kaare Mikkelsen , Preben Kidmose

A singing voice conversion model converts a song in the voice of an arbitrary source singer to the voice of a target singer. Recently, methods that leverage self-supervised audio representations such as HuBERT and Wav2Vec 2.0 have helped…

音频与语音处理 · 电气工程与系统科学 2023-03-23 Tejas Jayashankar , Jilong Wu , Leda Sari , David Kant , Vimal Manohar , Qing He