中文
相关论文

相关论文: Vocoder-Based Speech Synthesis from Silent Videos

200 篇论文

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

多媒体 · 计算机科学 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and acoustic scene…

音频与语音处理 · 电气工程与系统科学 2020-07-14 Abhinav Shukla , Stavros Petridis , Maja Pantic

In this paper, we propose a novel Lip-to-Speech synthesis (L2S) framework, for synthesizing intelligible speech from a silent lip movement video. Specifically, to complement the insufficient supervisory signal of the previous L2S model, we…

声音 · 计算机科学 2023-06-01 Jeongsoo Choi , Minsu Kim , Yong Man Ro

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis systems only consider the information within each sentence…

声音 · 计算机科学 2023-05-19 Ya-Jie Zhang , Wei Song , Yanghao Yue , Zhengchen Zhang , Youzheng Wu , Xiaodong He

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

The recent state of the art on monocular 3D face reconstruction from image data has made some impressive advancements, thanks to the advent of Deep Learning. However, it has mostly focused on input coming from a single RGB image,…

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

We introduce a deep learning model for speech denoising, a long-standing challenge in audio analysis arising in numerous applications. Our approach is based on a key observation about human speech: there is often a short pause between each…

声音 · 计算机科学 2020-10-26 Ruilin Xu , Rundi Wu , Yuko Ishiwaka , Carl Vondrick , Changxi Zheng

We present an unsupervised approach that converts the input speech of any individual into audiovisual streams of potentially-infinitely many output speakers. Our approach builds on simple autoencoders that project out-of-sample data onto…

计算机视觉与模式识别 · 计算机科学 2021-07-06 Kangle Deng , Aayush Bansal , Deva Ramanan

In this paper, we focus on improving the performance of the text-dependent speaker verification system in the scenario of limited training data. The speaker verification system deep learning based text-dependent generally needs a large…

声音 · 计算机科学 2020-11-24 Xiaoyi Qin , Yaogen Yang , Lin Yang , Xuyang Wang , Junjie Wang , Ming Li

We propose a semi-supervised singing synthesizer, which is able to learn new voices from audio data only, without any annotations such as phonetic segmentation. Our system is an encoder-decoder model with two encoders, linguistic and…

声音 · 计算机科学 2020-11-06 Jordi Bonada , Merlijn Blaauw

Despite the close relationship between speech perception and production, research in automatic speech recognition (ASR) and text-to-speech synthesis (TTS) has progressed more or less independently without exerting much mutual influence on…

计算与语言 · 计算机科学 2017-07-18 Andros Tjandra , Sakriani Sakti , Satoshi Nakamura

High-fidelity speech can be synthesized by end-to-end text-to-speech models in recent years. However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system remains a challenge.…

音频与语音处理 · 电气工程与系统科学 2020-08-05 Zexin Cai , Chuxiong Zhang , Ming Li

In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…

音频与语音处理 · 电气工程与系统科学 2021-06-17 Hyun Gon Ryu , Jeong-Hoon Kim , Simon See

In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extremely simple approach to generating (weak) speech…

多媒体 · 计算机科学 2017-06-02 Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , Ian Sturdy

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Karren Yang , Dejan Markovic , Steven Krenn , Vasu Agrawal , Alexander Richard

Audiovisual speech synthesis is the problem of synthesizing a talking face while maximizing the coherency of the acoustic and visual speech. In this paper, we propose and compare two audiovisual speech synthesis systems for 3D face models.…

音频与语音处理 · 电气工程与系统科学 2021-08-31 Ahmed Hussen Abdelaziz , Anushree Prasanna Kumar , Chloe Seivwright , Gabriele Fanelli , Justin Binder , Yannis Stylianou , Sachin Kajarekar

With the popularity of deep neural network, speech synthesis task has achieved significant improvements based on the end-to-end encoder-decoder framework in the recent days. More and more applications relying on speech synthesis technology…

音频与语音处理 · 电气工程与系统科学 2020-10-23 Dongyang Dai , Li Chen , Yuping Wang , Mu Wang , Rui Xia , Xuchen Song , Zhiyong Wu , Yuxuan Wang