English
Related papers

Related papers: LipFormer: Learning to Lipread Unseen Speakers bas…

200 papers

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Luchuan Song , Pinxin Liu , Haiyang Liu , Zhenchao Jin , Yolo Yunlong Tang , Zichong Xu , Susan Liang , Jing Bi , Jason J Corso , Chenliang Xu

In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to transcribe the audio as well as identify the speakers for downstream applications. Since overlapped speech is common in this…

Sound · Computer Science 2021-04-07 Liang Lu , Naoyuki Kanda , Jinyu Li , Yifan Gong

Despite the tremendous progress in zero-shot learning(ZSL), the majority of existing methods still rely on human-annotated attributes, which are difficult to annotate and scale. An unsupervised alternative is to represent each class using…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Muhammad Ferjad Naeem , Yongqin Xian , Luc Van Gool , Federico Tombari

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

Image and Video Processing · Electrical Eng. & Systems 2022-09-27 Rahul Sharma , Shrikanth Narayanan

This paper presents a novel approach for generating 3D talking heads from raw audio inputs. Our method grounds on the idea that speech related movements can be comprehensively and efficiently described by the motion of a few control points…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Federico Nocentini , Claudio Ferrari , Stefano Berretti

As a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in…

Multimedia · Computer Science 2021-12-21 Jinglin Liu , Zhiying Zhu , Yi Ren , Wencan Huang , Baoxing Huai , Nicholas Yuan , Zhou Zhao

Driven by deep learning techniques and large-scale datasets, recent years have witnessed a paradigm shift in automatic lip reading. While the main thrust of Visual Speech Recognition (VSR) was improving accuracy of Audio Speech Recognition…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Marzieh Oghbaie , Arian Sabaghi , Kooshan Hashemifard , Mohammad Akbari

Speech-driven 3D facial animation aims to synthesize vivid facial animations that accurately synchronize with speech and match the unique speaking style. However, existing works primarily focus on achieving precise lip synchronization while…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Hui Fu , Zeqing Wang , Ke Gong , Keze Wang , Tianshui Chen , Haojie Li , Haifeng Zeng , Wenxiong Kang

Although current deep learning-based face forgery detectors achieve impressive performance in constrained scenarios, they are vulnerable to samples created by unseen manipulation methods. Some recent works show improvements in…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Alexandros Haliassos , Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Audio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, the existing GAN-based models emphasize generating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Jintao Tan , Xize Cheng , Lingyu Xiong , Lei Zhu , Xiandong Li , Xianjia Wu , Kai Gong , Minglei Li , Yi Cai

Describing images with text is a fundamental problem in vision-language research. Current studies in this domain mostly focus on single image captioning. However, in various real applications (e.g., image editing, difference interpretation,…

Computation and Language · Computer Science 2019-06-20 Hao Tan , Franck Dernoncourt , Zhe Lin , Trung Bui , Mohit Bansal

When a multimodal Transformer answers a visual question, is the prediction driven by visual evidence, linguistic reasoning, or genuinely fused cross-modal computation -- and how does this structure evolve across layers? We address this…

Artificial Intelligence · Computer Science 2026-02-18 Hongxuan Wu , Yukun Zhang , Xueqing Zhou

Data-driven speech processing models usually perform well with a large amount of text supervision, but collecting transcribed speech data is costly. Therefore, we propose SpeechCLIP, a novel framework bridging speech and text through images…

Computation and Language · Computer Science 2022-10-26 Yi-Jen Shih , Hsuan-Fu Wang , Heng-Jui Chang , Layne Berry , Hung-yi Lee , David Harwath

In this work, we explore a new problem of frame interpolation for speech videos. Such content today forms the major form of online communication. We try to solve this problem by using several deep learning video generation algorithms to…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Aradhya Neeraj Mathur , Devansh Batra , Yaman Kumar , Rajiv Ratn Shah , Roger Zimmermann

This paper presents an efficient visual speech encoder for lip reading. While most recent lip reading studies have been based on the ResNet architecture and have achieved significant success, they are not sufficiently suitable for…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Young-Hu Park , Rae-Hong Park , Hyung-Min Park

Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continuous speech).…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-19 Helen L Bear

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

Non-frontal lip views contain useful information which can be used to enhance the performance of frontal view lipreading. However, the vast majority of recent lipreading works, including the deep learning approaches which significantly…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Stavros Petridis , Yujiang Wang , Zuwei Li , Maja Pantic

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Se Jin Park , Minsu Kim , Joanna Hong , Jeongsoo Choi , Yong Man Ro

We propose Masker, an unsupervised text-editing method for style transfer. To tackle cases when no parallel source-target pairs are available, we train masked language models (MLMs) for both the source and the target domain. Then we find…

Computation and Language · Computer Science 2020-10-05 Eric Malmi , Aliaksei Severyn , Sascha Rothe
‹ Prev 1 8 9 10 Next ›