中文
相关论文

相关论文: How Video Meetings Change Your Expression

200 篇论文

In this paper, we investigate the problem of unpaired video-to-video translation. Given a video in the source domain, we aim to learn the conditional distribution of the corresponding video in the target domain, without seeing any pairs of…

计算机视觉与模式识别 · 计算机科学 2019-08-22 Kwanyong Park , Sanghyun Woo , Dahun Kim , Donghyeon Cho , In So Kweon

Image-to-image translation and voice conversion enable the generation of a new facial image and voice while maintaining some of the semantics such as a pose in an image and linguistic content in audio, respectively. They can aid in the…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Naoya Takahashi , Mayank K. Singh , Yuki Mitsufuji

Creating realistic, natural, and lip-readable talking face videos remains a formidable challenge. Previous research primarily concentrated on generating and aligning single-frame images while overlooking the smoothness of frame-to-frame…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Shuheng Ge , Haoyu Xing , Li Zhang , Xiangqian Wu

Creating artificial social intelligence - algorithms that can understand the nuances of multi-person interactions - is an exciting and emerging challenge in processing facial expressions and gestures from multimodal videos. Recent…

机器学习 · 计算机科学 2022-10-28 Alex Wilf , Martin Q. Ma , Paul Pu Liang , Amir Zadeh , Louis-Philippe Morency

This paper studies an efficient multimodal data communication scheme for video conferencing. In our considered system, a speaker gives a talk to the audiences, with talking head video and audio being transmitted. Since the speaker does not…

多媒体 · 计算机科学 2024-10-30 Haonan Tong , Haopeng Li , Hongyang Du , Zhaohui Yang , Changchuan Yin , Dusit Niyato

End-to-End intelligent neural dialogue systems suffer from the problems of generating inconsistent and repetitive responses. Existing dialogue models pay attention to unilaterally incorporating personal knowledge into the dialog while…

计算与语言 · 计算机科学 2021-07-19 Yajing Sun , Yue Hu , Luxi Xing , Yuqiang Xie , Xiangpeng Wei

Human face-to-face conversation is an ideal model for human-computer dialogue. One of the major features of face-to-face communication is its multiplicity of communication channels that act on multiple modalities. To realize a natural…

cmp-lg · 计算机科学 2008-02-03 Katashi Nagao , Akikazu Takeuchi

Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few…

计算机视觉与模式识别 · 计算机科学 2023-04-21 Shuai Shen , Wenliang Zhao , Zibin Meng , Wanhua Li , Zheng Zhu , Jie Zhou , Jiwen Lu

What if the patterns hidden within dialogue reveal more about communication than the words themselves? We introduce Conversational DNA, a novel visual language that treats any dialogue -- whether between humans, between human and AI, or…

人机交互 · 计算机科学 2025-08-12 Baihan Lin

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Joon Son Chung , Shinji Watanabe

Video captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling object interactions.…

计算机视觉与模式识别 · 计算机科学 2020-04-01 Boxiao Pan , Haoye Cai , De-An Huang , Kuan-Hui Lee , Adrien Gaidon , Ehsan Adeli , Juan Carlos Niebles

Gaze is an essential prompt for analyzing human behavior and attention. Recently, there has been an increasing interest in determining gaze direction from facial videos. However, video gaze estimation faces significant challenges, such as…

计算机视觉与模式识别 · 计算机科学 2024-04-11 Swati Jindal , Mohit Yadav , Roberto Manduchi

How much can we infer about an emotional voice solely from an expressive face? This intriguing question holds great potential for applications such as virtual character dubbing and aiding individuals with expressive language disorders.…

声音 · 计算机科学 2025-02-04 Jiaxin Ye , Boyuan Cao , Hongming Shan

How much can we infer about a person's looks from the way they speak? In this paper, we study the task of reconstructing a facial image of a person from a short audio recording of that person speaking. We design and train a deep neural…

计算机视觉与模式识别 · 计算机科学 2019-05-24 Tae-Hyun Oh , Tali Dekel , Changil Kim , Inbar Mosseri , William T. Freeman , Michael Rubinstein , Wojciech Matusik

We present a versatile model, FaceAnime, for various video generation tasks from still images. Video generation from a single face image is an interesting problem and usually tackled by utilizing Generative Adversarial Networks (GANs) to…

计算机视觉与模式识别 · 计算机科学 2021-06-01 Xiaoguang Tu , Yingtian Zou , Jian Zhao , Wenjie Ai , Jian Dong , Yuan Yao , Zhikang Wang , Guodong Guo , Zhifeng Li , Wei Liu , Jiashi Feng

We propose a flexible probabilistic model for predicting turn-taking patterns in group conversations based solely on individual characteristics and past speaking behavior. Many models of conversation dynamics cannot yield insights that…

机器学习 · 计算机科学 2025-10-22 Madeline Navarro , Lisa O'Bryan , Santiago Segarra

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

计算机视觉与模式识别 · 计算机科学 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

Neural networks have recently become good at engaging in dialog. However, current approaches are based solely on verbal text, lacking the richness of a real face-to-face conversation. We propose a neural conversation model that aims to read…

计算机视觉与模式识别 · 计算机科学 2018-12-05 Hang Chu , Daiqing Li , Sanja Fidler

Given the audio-visual clip of the speaker, facial reaction generation aims to predict the listener's facial reactions. The challenge lies in capturing the relevance between video and audio while balancing appropriateness, realism, and…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Jiaming Li , Sheng Wang , Xin Wang , Yitao Zhu , Honglin Xiong , Zixu Zhuang , Qian Wang

Motivated by the previous success of Two-Dimensional Convolutional Neural Network (2D CNN) on image recognition, researchers endeavor to leverage it to characterize videos. However, one limitation of applying 2D CNN to analyze videos is…

计算机视觉与模式识别 · 计算机科学 2020-07-16 Junwu Weng , Donghao Luo , Yabiao Wang , Ying Tai , Chengjie Wang , Jilin Li , Feiyue Huang , Xudong Jiang , Junsong Yuan