English
Related papers

Related papers: ISExplore:Informative Segment Selection for Effici…

200 papers

In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead…

Graphics · Computer Science 2022-12-09 Zhentao Yu , Zixin Yin , Deyu Zhou , Duomin Wang , Finn Wong , Baoyuan Wang

In speech emotion recognition (SER), using predefined features without considering their practical importance may lead to high dimensional datasets, including redundant and irrelevant information. Consequently, high-dimensional learning…

Sound · Computer Science 2024-06-07 Alaa Nfissi , Wassim Bouachir , Nizar Bouguila , Brian Mishara

Generating long, high-quality videos remains a challenge due to the complex interplay of spatial and temporal dynamics and hardware limitations. In this work, we introduce MaskFlow, a unified video generation framework that combines…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Michael Fuest , Vincent Tao Hu , Björn Ommer

This paper discusses the task of face-based speech synthesis, a kind of personalized speech synthesis where the synthesized voices are constrained to perceptually match with a reference face image. Due to the lack of TTS-quality…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-07 Yao Shi , Yunfei Xu , Hongbin Suo , Yulong Wan , Haifeng Liu

Referring expression generation (REG) models that use speaker-dependent information require a considerable amount of training data produced by every individual speaker, or may otherwise perform poorly. In this work we present a simple REG…

Computation and Language · Computer Science 2017-04-13 Thiago castro Ferreira , Ivandre Paraboni

Learning performance data (e.g., quiz scores and attempts) is significant for understanding learner engagement and knowledge mastery level. However, the learning performance data collected from Intelligent Tutoring Systems (ITSs) often…

Computers and Society · Computer Science 2024-02-06 Liang Zhang , Jionghao Lin , Conrad Borchers , Meng Cao , Xiangen Hu

Effective human behavior modeling is critical for successful human-robot interaction. Current state-of-the-art approaches for predicting listening head behavior during dyadic conversations employ continuous-to-discrete representations,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Tri Tung Nguyen Nguyen , Quang Tien Dam , Dinh Tuan Tran , Joo-Ho Lee

We present Free-HeadGAN, a person-generic neural talking head synthesis system. We show that modeling faces with sparse 3D facial landmarks are sufficient for achieving state-of-the-art generative performance, without relying on strong…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Michail Christos Doukas , Evangelos Ververas , Viktoriia Sharmanska , Stefanos Zafeiriou

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Wenqing Wang , Yun Fu

In this work, we introduce a method that learns a single dynamic neural radiance field (NeRF) from monocular talking face videos of multiple identities. NeRFs have shown remarkable results in modeling the 4D dynamics and appearance of human…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Aggelina Chatziagapi , Grigorios G. Chrysos , Dimitris Samaras

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi

Video question-answering is a fundamental task in the field of video understanding. Although current vision--language models (VLMs) equipped with Video Transformers have enabled temporal modeling and yielded superior results, they are at…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Wei Han , Hui Chen , Min-Yen Kan , Soujanya Poria

Large language models (LLMs) demonstrate exceptional performance in numerous tasks but still heavily rely on knowledge stored in their parameters. Moreover, updating this knowledge incurs high training costs. Retrieval-augmented generation…

Computation and Language · Computer Science 2024-06-07 Yanming Liu , Xinyue Peng , Xuhong Zhang , Weihao Liu , Jianwei Yin , Jiannan Cao , Tianyu Du

In recent years, speech generation has seen remarkable progress, now achieving one-shot generation capability that is often virtually indistinguishable from real human voice. Integrating such advancements in speech generation with large…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-08 Paweł Budzianowski , Taras Sereda , Tomasz Cichy , Ivan Vulić

3D Gaussian splatting (3DGS) is an innovative rendering technique that surpasses the neural radiance field (NeRF) in both rendering speed and visual quality by leveraging an explicit 3D scene representation. Existing 3DGS approaches require…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Lintao Xiang , Hongpei Zheng , Yating Huang , Qijun Yang , Hujun Yin

Generalizable neural radiance field (NeRF) enables neural-based digital human rendering without per-scene retraining. When combined with human prior knowledge, high-quality human rendering can be achieved even with sparse input views.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Zhaorong Wang , Yoshihiro Kanamori , Yuki Endo

The growing data sharing and life-logging cultures are driving an unprecedented increase in the amount of unedited First-Person Videos. In this paper, we address the problem of accessing relevant information in First-Person Videos by…

Computer Vision and Pattern Recognition · Computer Science 2020-06-11 Alan Carvalho Neves , Michel Melo Silva , Mario Fernando Montenegro Campos , Erickson Rangel Nascimento

Neural Radiance Fields (NeRF) have recently emerged as a powerful method for image-based 3D reconstruction, but the lengthy per-scene optimization limits their practical usage, especially in resource-constrained settings. Existing…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Marco Orsingher , Anthony Dell'Eva , Paolo Zani , Paolo Medici , Massimo Bertozzi

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Suzhen Wang , Yifeng Ma , Yu Ding , Zhipeng Hu , Changjie Fan , Tangjie Lv , Zhidong Deng , Xin Yu

Despite numerous completed studies, achieving high fidelity talking face generation with highly synchronized lip movements corresponding to arbitrary audio remains a significant challenge in the field. The shortcomings of published studies…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Juan Zhang , Jiahao Chen , Cheng Wang , Zhiwang Yu , Tangquan Qi , Di Wu
‹ Prev 1 4 5 6 7 8 10 Next ›