中文
相关论文

相关论文: Vocoder-Based Speech Synthesis from Silent Videos

200 篇论文

Humans are able to imagine a person's voice from the person's appearance and imagine the person's appearance from his/her voice. In this paper, we make the first attempt to develop a method that can convert speech into a voice that matches…

声音 · 计算机科学 2019-04-10 Hirokazu Kameoka , Kou Tanaka , Aaron Valero Puche , Yasunori Ohishi , Takuhiro Kaneko

Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mainly use sequence-to-sequence models to solve this problem,…

声音 · 计算机科学 2022-07-14 Yongqi Wang , Zhou Zhao

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works…

计算与语言 · 计算机科学 2025-10-30 Shreyas Gopal , Ashutosh Anshul , Haoyang Li , Yue Heng Yeo , Hexin Liu , Eng Siong Chng

Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acquire. In this work, we approach text-to-audio synthesis using…

Deploying speech enhancement (SE) systems in wearable devices, such as smart glasses, is challenging due to the limited computational resources on the device. Although deep learning methods have achieved high-quality results, their…

音频与语音处理 · 电气工程与系统科学 2025-08-21 Heitor R. Guimarães , Ke Tan , Juan Azcarreta , Jesus Alvarez , Prabhav Agrawal , Ashutosh Pandey , Buye Xu

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

AI-synthesized speech, also known as deepfake speech, has recently raised significant concerns due to the rapid advancement of speech synthesis and speech conversion techniques. Previous works often rely on distinguishing synthesizer…

声音 · 计算机科学 2024-11-15 Kuiyuan Zhang , Zhongyun Hua , Yushu Zhang , Yifang Guo , Tao Xiang

A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these components often requires extensive domain expertise and may contain…

This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction…

声音 · 计算机科学 2021-07-15 Shijing Si , Jianzong Wang , Xiaoyang Qu , Ning Cheng , Wenqi Wei , Xinghua Zhu , Jing Xiao

Existing deep learning based speech enhancement mainly employ a data-driven approach, which leverage large amounts of data with a variety of noise types to achieve noise removal from noisy signal. However, the high dependence on the data…

声音 · 计算机科学 2024-01-24 Huaying Xue , Xiulian Peng , Yan Lu

In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the recent speech recognition and lip reading models are…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Devesh Walawalkar , Yihui He , Rohit Pillai

Conversational speech not only contains several variants of neutral speech but is also prominently interlaced with several speaker generated non-speech sounds such as laughter and breath. A robust speaker recognition system should be…

声音 · 计算机科学 2017-05-29 Sri Harsha Dumpala , Ashish Panda , Sunil Kumar Kopparapu

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority…

音频与语音处理 · 电气工程与系统科学 2022-02-22 Viet Anh Trinh , Sebastian Braun

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking…

Silent speech interfaces (SSI) has been an exciting area of recent interest. In this paper, we present a non-invasive silent speech interface that uses inaudible acoustic signals to capture people's lip movements when they speak. We exploit…

音频与语音处理 · 电气工程与系统科学 2020-11-24 Jian Luo , Jianzong Wang , Ning Cheng , Guilin Jiang , Jing Xiao

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

计算与语言 · 计算机科学 2025-04-11 Lakshmipathi Balaji , Karan Singla

Learning-based Text To Speech systems have the potential to generalize from one speaker to the next and thus require a relatively short sample of any new voice. However, this promise is currently largely unrealized. We present a method that…

机器学习 · 计算机科学 2018-02-21 Eliya Nachmani , Adam Polyak , Yaniv Taigman , Lior Wolf

Although text-to-speech (TTS) systems have significantly improved, most TTS systems still have limitations in synthesizing speech with appropriate phrasing. For natural speech synthesis, it is important to synthesize the speech with a…

音频与语音处理 · 电气工程与系统科学 2023-06-14 Ji-Sang Hwang , Sang-Hoon Lee , Seong-Whan Lee

Synthesizing realistic co-speech gestures is an important and yet unsolved problem for creating believable motions that can drive a humanoid robot to interact and communicate with human users. Such capability will improve the impressions of…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Shuhong Lu , Youngwoo Yoon , Andrew Feng

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li