English
Related papers

Related papers: EasyVoice: Integrating voice synthesis with Skype

200 papers

Zero-shot voice conversion aims to transfer the voice of a source speaker to that of a speaker unseen during training, while preserving the content information. Although various methods have been proposed to reconstruct speaker information…

Sound · Computer Science 2024-08-22 Anastasia Avdeeva , Aleksei Gusev

The goal of voice conversion (VC) is to convert input voice to match the target speaker's voice while keeping text and prosody intact. VC is usually used in entertainment and speaking-aid systems, as well as applied for speech data…

Sound · Computer Science 2022-04-01 A. Kashkin , I. Karpukhin , S. Shishkin

Modern text-to-speech systems are able to produce natural and high-quality speech, but speech contains factors of variation (e.g. pitch, rhythm, loudness, timbre)\ that text alone cannot contain. In this work we move towards a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-29 Giorgio Fabbro , Vladimir Golkov , Thomas Kemp , Daniel Cremers

The inclusion of voice persona in synthesized voice can be significant in a broad range of human-computer-interaction (HCI) applications, including augmentative and assistive communication (AAC), artistic performance, and design of virtual…

Sound · Computer Science 2022-11-01 Camille Noufi , Lloyd May , Jonathan Berger

It has always been a rather tough task to communicate with someone possessing a hearing impairment. One of the most tested ways to establish such a communication is through the use of sign based languages. However, not many people are aware…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Sharanya Mukherjee , Md Hishaam Akhtar , Kannadasan R

Haptic feedback is the most significant sensory interface following visual cues. Developing thin, flexible surfaces that function as haptic interfaces is important for augmenting virtual reality, wearable devices, robotics and prostheses.…

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world.…

Sound · Computer Science 2026-02-04 Chengyuan Ma , Jiawei Jin , Ruijie Xiong , Chunxiang Jin , Canxiang Yan , Wenming Yang

It is challenging to build a multi-singer high-fidelity singing voice synthesis system with cross-lingual ability by only using monolingual singers in the training stage. In this paper, we propose CrossSinger, which is a cross-lingual…

Sound · Computer Science 2023-09-25 Xintong Wang , Chang Zeng , Jun Chen , Chunhui Wang

We propose a singing decomposition system that encodes time-aligned linguistic content, pitch, and source speaker identity via Assem-VC. With decomposed speaker-independent information and the target speaker's embedding, we could synthesize…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-26 Kang-wook Kim , Junhyeok Lee

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

Sound plays a significant role in human memory, yet it is often overlooked by mainstream life-recording methods. Most current UGC (User-Generated Content) creation tools emphasize visual content while lacking user-friendly sound design…

Human-Computer Interaction · Computer Science 2024-10-11 Chongjun Zhong , Jiaxing Yu , Yingping Cao , Songruoyao Wu , Wenqi Wu , Kejun Zhang

An integral part of spreadsheet auditing is navigation. For sufferers of Repetitive Strain Injury who need to use voice recognition technology this navigation can be highly problematic. To counter this the authors have developed an…

Human-Computer Interaction · Computer Science 2008-09-23 Derek Flood , Kevin Mc Daid , Fergal Mc Caffery , Brian Bishop

Synthesized speech is common today due to the prevalence of virtual assistants, easy-to-use tools for generating and modifying speech signals, and remote work practices. Synthesized speech can also be used for nefarious purposes, including…

Sound · Computer Science 2022-05-05 Emily R. Bartusiak , Edward J. Delp

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary annotations, limiting…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-22 Yu Zhang , Wenxiang Guo , Changhao Pan , Dongyu Yao , Zhiyuan Zhu , Ziyue Jiang , Yuhan Wang , Tao Jin , Zhou Zhao

This paper proposes a controllable singing voice synthesis system capable of generating expressive singing voice with two novel methodologies. First, a local style token module, which predicts frame-level style tokens from an input pitch…

Sound · Computer Science 2022-04-08 Juheon Lee , Hyeong-Seok Choi , Kyogu Lee

Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper proposes to unify these…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-23 Wei-Ning Hsu , Tal Remez , Bowen Shi , Jacob Donley , Yossi Adi

This paper aims to introduce a robust singing voice synthesis (SVS) system to produce very natural and realistic singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random…

Sound · Computer Science 2023-02-17 Zewang Zhang , Yibin Zheng , Xinhui Li , Li Lu

Removing background noise from speech audio has been the subject of considerable effort, especially in recent years due to the rise of virtual communication and amateur recordings. Yet background noise is not the only unpleasant disturbance…

Sound · Computer Science 2022-09-19 Joan Serrà , Santiago Pascual , Jordi Pons , R. Oguz Araz , Davide Scaini

In this paper, a text-to-rapping/singing system is introduced, which can be adapted to any speaker's voice. It utilizes a Tacotron-based multispeaker acoustic model trained on read-only speech data and which provides prosody control at the…

Diffusion models have demonstrated remarkable performance in speech synthesis, but typically require multi-step sampling, resulting in low inference efficiency. Recent studies address this issue by distilling diffusion models into…

Sound · Computer Science 2025-10-08 Tao Zhu , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng