中文
相关论文

相关论文: I'm Sorry for Your Loss: Spectrally-Based Audio Di…

200 篇论文

Anomaly detection has many important applications, such as monitoring industrial equipment. Despite recent advances in anomaly detection with deep-learning methods, it is unclear how existing solutions would perform under…

声音 · 计算机科学 2022-04-06 Bingqing Chen , Luca Bondi , Samarjit Das

The performance of voice-based Parkinson's disease (PD) detection systems degrades when there is an acoustic mismatch between training and operating conditions caused mainly by degradation in test signals. In this paper, we address this…

Nowadays, we have witnessed the early progress on learning the association between voice and face automatically, which brings a new wave of studies to the computer vision community. However, most of the prior arts along this line (a) merely…

计算机视觉与模式识别 · 计算机科学 2021-03-15 Peisong Wen , Qianqian Xu , Yangbangyan Jiang , Zhiyong Yang , Yuan He , Qingming Huang

Few-shot learning has emerged as a powerful paradigm for training models with limited labeled data, addressing challenges in scenarios where large-scale annotation is impractical. While extensive research has been conducted in the image…

Speech audio quality is subject to degradation caused by an acoustic environment and isotropic ambient and point noises. The environment can lead to decreased speech intelligibility and loss of focus and attention by the listener. Basic…

音频与语音处理 · 电气工程与系统科学 2022-04-05 Paula Sánchez López , Paul Callens , Milos Cernak

Music can be represented in multiple forms, such as in the audio form as a recording of a performance, in the symbolic form as a computer readable score, or in the image form as a scan of the sheet music. Music synchronisation provides a…

声音 · 计算机科学 2022-06-02 Ruchit Agrawal

The field of prosody transfer in speech synthesis systems is rapidly advancing. This research is focused on evaluating learning methods for adapting pre-trained monolingual text-to-speech (TTS) models to multilingual conditions, i.e.,…

计算与语言 · 计算机科学 2024-06-19 Arnav Goel , Medha Hira , Anubha Gupta

We consider a wireless sensor network, sampling a bandlimited field, described by a limited number of harmonics. Sensor nodes are irregularly deployed over the area of interest or subject to random motion; in addition sensors measurements…

其他计算机科学 · 计算机科学 2009-11-13 A. Nordio , C. -F. Chiasserini , E. Viterbo

Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal…

音频与语音处理 · 电气工程与系统科学 2025-06-10 Aníbal J. S. Ferreira , Luis M. T. Jesus , Laurentino M. M. Leal , Jorge E. F. Spratley

Suffering from limited singing voice corpus, existing singing voice synthesis (SVS) methods that build encoder-decoder neural networks to directly generate spectrogram could lead to out-of-tune issues during the inference phase. To…

声音 · 计算机科学 2021-10-13 Shujun Liu , Hai Zhu , Kun Wang , Huajun Wang

While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This would be highly desirable, as it adds explainability to the…

音频与语音处理 · 电气工程与系统科学 2025-10-10 Michael Kuhlmann , Fritz Seebauer , Petra Wagner , Reinhold Haeb-Umbach

Automatic speech quality assessment is an important, transversal task whose progress is hampered by the scarcity of human annotations, poor generalization to unseen recording conditions, and a lack of flexibility of existing approaches. In…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Joan Serrà , Jordi Pons , Santiago Pascual

With the advent of modern AI architectures, a shift has happened towards end-to-end architectures. This pivot has led to neural architectures being trained without domain-specific biases/knowledge, optimized according to the task. We in…

声音 · 计算机科学 2025-05-08 Prateek Verma

Sound synthesis is a complex field that requires domain expertise. Manual tuning of synthesizer parameters to match a specific sound can be an exhaustive task, even for experienced sound engineers. In this paper, we introduce InverSynth -…

声音 · 计算机科学 2019-11-22 Oren Barkan , David Tsiris , Ori Katz , Noam Koenigstein

In a spoken dialogue system, an NLU model is preceded by a speech recognition system that can deteriorate the performance of natural language understanding. This paper proposes a method for investigating the impact of speech recognition…

计算与语言 · 计算机科学 2023-10-26 Marek Kubis , Paweł Skórzewski , Marcin Sowański , Tomasz Ziętkiewicz

The ability to localize and track acoustic events is a fundamental prerequisite for equipping machines with the ability to be aware of and engage with humans in their surrounding environment. However, in realistic scenarios, audio signals…

音频与语音处理 · 电气工程与系统科学 2020-10-22 Christine Evers , Heinrich Loellmann , Heinrich Mellmann , Alexander Schmidt , Hendrik Barfuss , Patrick Naylor , Walter Kellermann

One of the primary sources of suboptimal image quality in ultrasound imaging is phase aberration. It is caused by spatial changes in sound speed over a heterogeneous medium, which disturbs the transmitted waves and prevents coherent…

图像与视频处理 · 电气工程与系统科学 2024-07-03 Mostafa Sharifzadeh , Sobhan Goudarzi , An Tang , Habib Benali , Hassan Rivaz

The correlation between the sharpness of loss minima and generalisation in the context of deep neural networks has been subject to discussion for a long time. Whilst mostly investigated in the context of selected benchmark data sets in the…

The identification of structural differences between a music performance and the score is a challenging yet integral step of audio-to-score alignment, an important subtask of music information retrieval. We present a novel method to detect…

声音 · 计算机科学 2021-02-16 Ruchit Agrawal , Daniel Wolff , Simon Dixon

Stuttering -- characterized by involuntary disfluencies such as blocks, prolongations, and repetitions -- is often misinterpreted by automatic speech recognition (ASR) systems, resulting in elevated word error rates and making voice-driven…

声音 · 计算机科学 2025-08-22 Dena Mujtaba , Nihar Mahapatra