English
Related papers

Related papers: Improving Inference-Time Optimisation for Vocal Ef…

200 papers

Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhaoxing Zhang , Junda Cheng , Gangwei Xu , Xiaoxiang Wang , Can Zhang , Xin Yang

Gaussian processes regression is applied to augment experimental data of transfer-path analysis (TPA) by known information about the underlying physical properties of the system under investigation. The approach can be used as an…

Classical Physics · Physics 2019-05-21 Christopher Albert

The goal of this work is zero-shot text-to-speech synthesis, with speaking styles and voices learnt from facial characteristics. Inspired by the natural fact that people can imagine the voice of someone when they look at his or her face, we…

Machine Learning · Computer Science 2023-02-28 Jiyoung Lee , Joon Son Chung , Soo-Whan Chung

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

Estimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60),…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Hanyu Meng , Jeroen Breebaart , Jeremy Stoddard , Vidhyasaharan Sethu , Eliathamby Ambikairajah

Despite rapid progress in the voice style transfer (VST) field, recent zero-shot VST systems still lack the ability to transfer the voice style of a novel speaker. In this paper, we present HierVST, a hierarchical adaptive end-to-end…

Sound · Computer Science 2023-08-01 Sang-Hoon Lee , Ha-Yeong Choi , Hyung-Seok Oh , Seong-Whan Lee

Singing voice beat tracking is a challenging task, due to the lack of musical accompaniment that often contains robust rhythmic and harmonic patterns, something most existing beat tracking systems utilize and can be essential for estimating…

Sound · Computer Science 2025-03-14 Jiajun Deng , Yaolong Ju , Jing Yang , Simon Lui , Xunying Liu

End-to-end model, especially Recurrent Neural Network Transducer (RNN-T), has achieved great success in speech recognition. However, transducer requires a great memory footprint and computing time when processing a long decoding sequence.…

Sound · Computer Science 2023-07-18 Xiaohui Zhang , Mangui Liang , Zhengkun Tian , Jiangyan Yi , Jianhua Tao

Diffusion models produce high-fidelity speech but are inefficient for real-time use due to long denoising steps and challenges in modeling intonation and rhythm. To improve this, we propose Diffusion Loss-Guided Policy Optimization (DLPO),…

Sound · Computer Science 2025-08-06 Jingyi Chen , Ju Seung Byun , Micha Elsner , Pichao Wang , Andrew Perrault

Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning systems. Prevailing learning paradigms have been relying on parallel audio-text data, which is, however,…

Sound · Computer Science 2022-05-04 Yanpeng Zhao , Jack Hessel , Youngjae Yu , Ximing Lu , Rowan Zellers , Yejin Choi

Image Style Transfer (IST) is an interdisciplinary topic of computer vision and art that continuously attracts researchers' interests. Different from traditional Image-guided Image Style Transfer (IIST) methods that require a style…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Hanyu Wang , Pengxiang Wu , Kevin Dela Rosa , Chen Wang , Abhinav Shrivastava

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models…

Audio-to-score alignment is an important pre-processing step for in-depth analysis of classical music. In this paper, we apply novel transposition-invariant audio features to this task. These low-dimensional features represent local pitch…

Sound · Computer Science 2018-07-20 Andreas Arzt , Stefan Lattner

Text Style Transfer (TST) aims to alter the underlying style of the source text to another specific style while keeping the same content. Due to the scarcity of high-quality parallel training data, unsupervised learning has become a…

Computation and Language · Computer Science 2021-12-07 Haoran Xu , Sixing Lu , Zhongkai Sun , Chengyuan Ma , Chenlei Guo

Speaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embeddings using deep neural networks for SV systems has gone…

Sound · Computer Science 2022-05-27 Nan Zhang , Jianzong Wang , Zhenhou Hong , Chendong Zhao , Xiaoyang Qu , Jing Xiao

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

Sound · Computer Science 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary domains, where either or both modalities can be distributionally shifted, which hampers online…

Machine Learning · Computer Science 2026-02-24 Sarthak Kumar Maharana , Akshay Mehra , Bhavya Ramakrishna , Yunhui Guo , Guan-Ming Su

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech. This…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-29 Zhichao Wang , Xinyong Zhou , Fengyu Yang , Tao Li , Hongqiang Du , Lei Xie , Wendong Gan , Haitao Chen , Hai Li

Sequence model learning algorithms typically maximize log-likelihood minus the norm of the model (or minimize Hamming loss + norm). In cross-lingual part-of-speech (POS) tagging, our target language training data consists of sequences of…

Computation and Language · Computer Science 2016-01-12 Anders Søgaard

Recent studies on inverse problems have proposed posterior samplers that leverage the pre-trained diffusion models as powerful priors. These attempts have paved the way for using diffusion models in a wide range of inverse problems.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Sojin Lee , Dogyun Park , Inho Kong , Hyunwoo J. Kim