English
Related papers

Related papers: FlexLip: A Controllable Text-to-Lip System

200 papers

Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those tasks into a SpeechLLM promises to exploit paralinguistic information in the speech and to…

Computation and Language · Computer Science 2026-05-15 Titouan Parcollet , Shucong Zhang , Xianrui Zheng , Rogier C. van Dalen

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Minsu Kim , Pingchuan Ma , Honglie Chen , Stavros Petridis , Maja Pantic

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Bowen Hao , Dongliang Zhou , Xiaojie Li , Xingyu Zhang , Liang Xie , Jianlong Wu , Erwei Yin

Lip synchronization aims to generate realistic talking videos that match given audio, which is essential for high-quality video dubbing. However, current methods have fundamental drawbacks: mask-based approaches suffer from local color…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Ruidi Fan , Yang Zhou , Siyuan Wang , Tian Yu , Yutong Jiang , Xusheng Liu

With the impressive progress in diffusion-based text-to-image generation, extending such powerful generative ability to text-to-video raises enormous attention. Existing methods either require large-scale text-video pairs and a large number…

Computer Vision and Pattern Recognition · Computer Science 2023-10-18 Ruiqi Wu , Liangyu Chen , Tong Yang , Chunle Guo , Chongyi Li , Xiangyu Zhang

Light-based advanced manufacturing increasingly requires programmable, closed-loop tools that translate human design intent into executable operations at small length scales. Yet a key bottleneck persists across robotic and manufacturing…

Robotics · Computer Science 2026-05-28 Ivan Saraev , Elena Erben , Weida Liao , Fan Nan , Gerhard Neumann , Eric Lauga , Moritz Kreysing

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Triantafyllos Afouras , Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

The goal of this work is to recognise phrases and sentences being spoken by a talking face, with or without the audio. Unlike previous works that have focussed on recognising a limited number of words or phrases, we tackle lip reading as an…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Joon Son Chung , Andrew Senior , Oriol Vinyals , Andrew Zisserman

Whenever we speak, our voice is accompanied by facial movements and expressions. Several recent works have shown the synthesis of highly photo-realistic videos of talking faces, but they either require a source video to drive the target…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Prateek Manocha , Prithwijit Guha

Recent image tone adjustment (or enhancement) approaches have predominantly adopted supervised learning for learning human-centric perceptual assessment. However, these approaches are constrained by intrinsic challenges of supervised…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Hyeongmin Lee , Kyoungkook Kang , Jungseul Ok , Sunghyun Cho

We propose ClipFace, a novel self-supervised approach for text-guided editing of textured 3D morphable model of faces. Specifically, we employ user-friendly language prompts to enable control of the expressions as well as appearance of 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Text-to-image generation models represent the next step of evolution in image synthesis, offering a natural way to achieve flexible yet fine-grained control over the result. One emerging area of research is the fast adaptation of large…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Anton Voronov , Mikhail Khoroshikh , Artem Babenko , Max Ryabinin

Audio is the primary modality for human communication and has driven the success of Automatic Speech Recognition (ASR) technologies. However, such audio-centric systems inherently exclude individuals who are deaf or hard of hearing. Visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jeong Hun Yeo , Hyeongseop Rha , Sungjune Park , Junil Won , Yong Man Ro

We propose an end-to-end lecture video generation system that can generate realistic and complete lecture videos directly from annotated slides, instructor's reference voice and instructor's reference portrait video. Our system is primarily…

Multimedia · Computer Science 2022-09-20 Wenbin Wang , Yang Song , Sanjay Jha

Taking inspiration from recent developments in visual generative tasks using diffusion models, we propose a method for end-to-end speech-driven video editing using a denoising diffusion model. Given a video of a talking person, and a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Dan Bigioi , Shubhajit Basak , Michał Stypułkowski , Maciej Zięba , Hugh Jordan , Rachel McDonnell , Peter Corcoran

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-08 Noé Tits

Editing talking-head video to change the speech content or to remove filler words is challenging. We propose a novel method to edit talking-head video based on its transcript to produce a realistic output video in which the dialogue of the…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Ohad Fried , Ayush Tewari , Michael Zollhöfer , Adam Finkelstein , Eli Shechtman , Dan B Goldman , Kyle Genova , Zeyu Jin , Christian Theobalt , Maneesh Agrawala

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edit such implicit…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Yuqi Sun , Ruian He , Weimin Tan , Bo Yan

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu
‹ Prev 1 8 9 10 Next ›