English
Related papers

Related papers: Beyond Monologue: Interactive Talking-Listening Av…

200 papers

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those methods struggle to learn…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Suzhen Wang , Lincheng Li , Yu Ding , Xin Yu

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond…

Artificial Intelligence · Computer Science 2025-05-06 Yemin Shi , Yu Shu , Siwei Dong , Guangyi Liu , Jaward Sesay , Jingwen Li , Zhiting Hu

Recently, large language models have facilitated the emergence of highly intelligent conversational AI capable of engaging in human-like dialogues. However, a notable distinction lies in the fact that these AI models predominantly generate…

Human-Computer Interaction · Computer Science 2025-10-13 Jijie Zhou , Yuhan Hu

This paper addresses the problem of generating lifelike holistic co-speech motions for 3D avatars, focusing on two key aspects: variability and coordination. Variability allows the avatar to exhibit a wide range of motions even with similar…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Yifei Liu , Qiong Cao , Yandong Wen , Huaiguang Jiang , Changxing Ding

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo

Codec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are almost indistinguishable from video. In this paper we describe…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Alexander Richard , Colin Lea , Shugao Ma , Juergen Gall , Fernando de la Torre , Yaser Sheikh

We present a large, tunable neural conversational response generation model, DialoGPT (dialogue generative pre-trained transformer). Trained on 147M conversation-like exchanges extracted from Reddit comment chains over a period spanning…

Computation and Language · Computer Science 2020-05-05 Yizhe Zhang , Siqi Sun , Michel Galley , Yen-Chun Chen , Chris Brockett , Xiang Gao , Jianfeng Gao , Jingjing Liu , Bill Dolan

Tuning language models for dialogue generation has been a prevalent paradigm for building capable dialogue agents. Yet, traditional tuning narrowly views dialogue generation as resembling other language generation tasks, ignoring the role…

Computation and Language · Computer Science 2024-05-31 Jian Wang , Chak Tou Leong , Jiashuo Wang , Dongding Lin , Wenjie Li , Xiao-Yong Wei

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Chenxu Zhang , Zenan Li , Hongyi Xu , You Xie , Xiaochen Zhao , Tianpei Gu , Guoxian Song , Xin Chen , Chao Liang , Jianwen Jiang , Linjie Luo

In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is to extend the field of co-speech gesture generation by…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Anna Deichler , Jim O'Regan , Jonas Beskow

While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Hang Zhou , Yasheng Sun , Wayne Wu , Chen Change Loy , Xiaogang Wang , Ziwei Liu

Recent advances in duplex speech models have enabled natural, low-latency speech-to-speech interactions. However, existing models are restricted to a fixed role and voice, limiting their ability to support structured, role-driven real-world…

Computation and Language · Computer Science 2026-02-09 Rajarshi Roy , Jonathan Raiman , Sang-gil Lee , Teodor-Dumitru Ene , Robert Kirby , Sungwon Kim , Jaehyeon Kim , Bryan Catanzaro

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Geumbyeol Hwang , Sunwon Hong , Seunghyun Lee , Sungwoo Park , Gyeongsu Chae

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Non-goal oriented dialog agents (i.e. chatbots) aim to produce varying and engaging conversations with a user; however, they typically exhibit either inconsistent personality across conversations or the average personality of all users.…

Computation and Language · Computer Science 2020-05-14 Alex Boyd , Raul Puri , Mohammad Shoeybi , Mostofa Patwary , Bryan Catanzaro

Conversational scenarios are very common in real-world settings, yet existing co-speech motion synthesis approaches often fall short in these contexts, where one person's audio and gestures will influence the other's responses.…

Sound · Computer Science 2024-12-04 Mingyi Shi , Dafei Qin , Leo Ho , Zhouyingcheng Liao , Yinghao Huang , Junichi Yamagishi , Taku Komura

Large Language Model (LLM)-enhanced agents become increasingly prevalent in Human-AI communication, offering vast potential from entertainment to professional domains. However, current multi-modal dialogue systems overlook the acoustic…

Computation and Language · Computer Science 2024-06-19 Haoqiu Yan , Yongxin Zhu , Kai Zheng , Bing Liu , Haoyu Cao , Deqiang Jiang , Linli Xu

In this paper, we present a methodology for the development of embodied conversational agents for social virtual worlds. The agents provide multimodal communication with their users in which speech interaction is included. Our proposal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-29 D. Griol , A. Sanchis , J. M. Molina , Z. Callejas
‹ Prev 1 4 5 6 7 8 10 Next ›