中文
相关论文

相关论文: VisualTTS: TTS with Accurate Lip-Speech Synchroniz…

200 篇论文

The goal of Automatic Voice Over (AVO) is to generate speech in sync with a silent video given its text script. Recent AVO frameworks built upon text-to-speech synthesis (TTS) have shown impressive results. However, the current AVO learning…

音频与语音处理 · 电气工程与系统科学 2023-06-30 Junchen Lu , Berrak Sisman , Mingyang Zhang , Haizhou Li

Dynamically synthesizing talking speech that actively responds to a listening head is critical during the face-to-face interaction. For example, the speaker could take advantage of the listener's facial expression to adjust the tones,…

音频与语音处理 · 电气工程与系统科学 2023-06-22 Mohan Zhou , Yalong Bai , Wei Zhang , Ting Yao , Tiejun Zhao , Tao Mei

The task of Visual Text-to-Speech (VisualTTS), also known as video dubbing, aims to generate speech synchronized with the lip movements in an input video, in additional to being consistent with the content of input text and cloning the…

多媒体 · 计算机科学 2025-12-01 Yuyue Wang , Xin Cheng , Yihan Wu , Xihua Wang , Jinchuan Tian , Ruihua Song

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual datasets, such as GRID,…

声音 · 计算机科学 2023-03-02 Zhe Niu , Brian Mak

In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos…

多媒体 · 计算机科学 2024-03-05 Sindhu Hegde , Rudrabha Mukhopadhyay , C. V. Jawahar , Vinay Namboodiri

This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into discrete symbols and synthesizes a speech waveform from…

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the…

In this paper, we address the problem of lip-voice synchronisation in videos containing human face and voice. Our approach is based on determining if the lips motion and the voice in a video are synchronised or not, depending on their…

计算机视觉与模式识别 · 计算机科学 2022-07-01 Venkatesh S. Kadandale , Juan F. Montesinos , Gloria Haro

Recently, artificial intelligence-based dubbing technology has advanced, enabling automated dubbing (AD) to convert the source speech of a video into target speech in different languages. However, natural AD still faces synchronization…

音频与语音处理 · 电气工程与系统科学 2026-05-05 Changi Hong , Yoonah Song , Hwayoung Park , Chaewoon Bang , Dayeon Ku , Do Hyun Lee , Hong Kook Kim

In this paper, we propose a neural end-to-end system for voice preserving, lip-synchronous translation of videos. The system is designed to combine multiple component models and produces a video of the original speaker speaking in the…

Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalness and audio-visual synchronization, and are limited to…

音频与语音处理 · 电气工程与系统科学 2025-12-08 Kaidi Wang , Yi He , Wenhao Guan , Weijie Wu , Hongwu Ding , Xiong Zhang , Di Wu , Meng Meng , Jian Luan , Lin Li , Qingyang Hong

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all…

音频与语音处理 · 电气工程与系统科学 2022-02-21 Disong Wang , Shan Yang , Dan Su , Xunying Liu , Dong Yu , Helen Meng

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

计算机视觉与模式识别 · 计算机科学 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

计算机视觉与模式识别 · 计算机科学 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

Traditional voice conversion (VC) methods typically attempt to separate speaker identity and linguistic information into distinct representations, which are then combined to reconstruct the audio. However, effectively disentangling these…

声音 · 计算机科学 2025-10-13 Huu Tuong Tu , Huan Vu , cuong tien nguyen , Dien Hy Ngo , Nguyen Thi Thu Trang

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is especially challenging…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Marshall Thomas , Edward Fish , Richard Bowden

Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip movements, VSR tasks require effectively utilizing any…

声音 · 计算机科学 2024-10-23 Zehua Liu , Xiaolou Li , Chen Chen , Li Guo , Lantian Li , Dong Wang

We study the problem of syncing the lip movement in a video with the audio stream. Our solution finds an optimal alignment using a dual-domain recurrent neural network that is trained on synthetic data we generate by dropping and…

计算机视觉与模式识别 · 计算机科学 2022-03-31 Yoav Shalev , Lior Wolf

Creating realistic, natural, and lip-readable talking face videos remains a formidable challenge. Previous research primarily concentrated on generating and aligning single-frame images while overlooking the smoothness of frame-to-frame…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Shuheng Ge , Haoyu Xing , Li Zhang , Xiangqian Wu

The recent text-to-speech (TTS) has achieved quality comparable to that of humans; however, its application in spoken dialogue has not been widely studied. This study aims to realize a TTS that closely resembles human dialogue. First, we…

音频与语音处理 · 电气工程与系统科学 2022-06-27 Kentaro Mitsui , Tianyu Zhao , Kei Sawada , Yukiya Hono , Yoshihiko Nankaku , Keiichi Tokuda
‹ 上一页 1 2 3 10 下一页 ›