English
Related papers

Related papers: MM-Conv: A Multi-modal Conversational Dataset for …

200 papers

In this work, we propose a framework that creates a lively virtual dynamic scene with contextual motions of multiple humans. Generating multi-human contextual motion requires holistic reasoning over dynamic relationships among human-human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Donggeun Lim , Jinseok Bae , Inwoo Hwang , Seungmin Lee , Hwanhee Lee , Young Min Kim

Embodied agents, in the form of virtual agents or social robots, are rapidly becoming more widespread. In human-human interactions, humans use nonverbal behaviours to convey their attitudes, feelings, and intentions. Therefore, this…

Artificial Intelligence · Computer Science 2026-04-30 Carson Yu Liu , Gelareh Mohammadi , Yang Song , Wafa Johal

Generating realistic conversational gestures are essential for achieving natural, socially engaging interactions with digital humans. However, existing methods typically map a single audio stream to a single speaker's motion, without…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yichen Peng , Jyun-Ting Song , Siyeol Jung , Ruofan Liu , Haiyang Liu , Xuangeng Chu , Ruicong Liu , Erwin Wu , Hideki Koike , Kris Kitani

Human-Machine Interaction (HMI) systems have gained huge interest in recent years, with reference expression comprehension being one of the main challenges. Traditionally human-machine interaction has been mostly limited to speech and…

Human-Computer Interaction · Computer Science 2023-06-21 Aman Jain , Anirudh Reddy Kondapally , Kentaro Yamada , Hitomi Yanaka

In the field of natural language processing, open-domain chatbots have emerged as an important research topic. However, a major limitation of existing open-domain chatbot research is its singular focus on short single-session dialogue,…

Computation and Language · Computer Science 2023-10-23 Jihyoung Jang , Minseong Boo , Hyounghun Kim

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

This paper studies an efficient multimodal data communication scheme for video conferencing. In our considered system, a speaker gives a talk to the audiences, with talking head video and audio being transmitted. Since the speaker does not…

Multimedia · Computer Science 2024-10-30 Haonan Tong , Haopeng Li , Hongyang Du , Zhaohui Yang , Changchuan Yin , Dusit Niyato

We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate in long and…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Mohan Zhou , Yalong Bai , Wei Zhang , Ting Yao , Tiejun Zhao

In this paper, we introduce ConversaSynth, a framework designed to generate synthetic conversation audio using large language models (LLMs) with multiple persona settings. The framework first creates diverse and coherent text-based…

Sound · Computer Science 2025-07-08 Kaung Myat Kyaw , Jonathan Hoyin Chan

We introduce Toyteller, an AI-powered storytelling system where users generate a mix of story text and visuals by directly manipulating character symbols like they are toy-playing. Anthropomorphized symbol motions can convey rich and…

Human-Computer Interaction · Computer Science 2025-01-24 John Joon Young Chung , Melissa Roemmele , Max Kreminski

Humans with an average level of social cognition can infer the beliefs of others based solely on the nonverbal communication signals (e.g. gaze, gesture, pose and contextual information) exhibited during social interactions. This social…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Jiafei Duan , Samson Yu , Nicholas Tan , Li Yi , Cheston Tan

In this study, conversations between humans and avatars are linguistically, organizationally, and structurally analyzed, focusing on what is necessary for creating face-to-face multimodal interfaces for machines. We videorecorded…

Human-Computer Interaction · Computer Science 2022-11-28 João Ranhel , Cacilda Vilela de Lima

We present SpeakingFaces as a publicly-available large-scale multimodal dataset developed to support machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include…

Human-Computer Interaction · Computer Science 2021-05-04 Madina Abdrakhmanova , Askat Kuzdeuov , Sheikh Jarju , Yerbolat Khassanov , Michael Lewis , Huseyin Atakan Varol

The capacity to create realistic virtual humans has progressed significantly, and such characters can be found in many applications across entertainment, education and health. As an essential element of interactive virtual humans,…

Graphics · Computer Science 2026-05-12 Haoyang Du , Yinghan Xu , John Dingliana , Brian Keegan , Rachel McDonnell , Cathy Ennis

When designing robots to assist in everyday human activities, it is crucial to enhance user requests with visual cues from their surroundings for improved intent understanding. This process is defined as a multimodal classification task.…

Computation and Language · Computer Science 2025-06-18 Shang-Chi Tsai , Seiya Kawano , Angel Garcia Contreras , Koichiro Yoshino , Yun-Nung Chen

Memes are widely used in online social interactions, providing vivid, intuitive, and often humorous means to express intentions and emotions. Existing dialogue datasets are predominantly limited to either manually annotated or pure-text…

Computation and Language · Computer Science 2025-07-02 Yuheng Wang , Xianhe Tang , Pufeng Huang

Real-time magnetic resonance imaging (RT-MRI) of human speech production is enabling significant advances in speech science, linguistics, bio-inspired speech technology development, and clinical applications. Easy access to RT-MRI is…

This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi-view captures of a subject, we create an animatable 3D face…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Richard Shaw , Youngkyoon Jang , Athanasios Papaioannou , Arthur Moreau , Helisa Dhamo , Zhensong Zhang , Eduardo Pérez-Pellitero

Recording the dynamics of unscripted human interactions in the wild is challenging due to the delicate trade-offs between several factors: participant privacy, ecological validity, data fidelity, and logistical overheads. To address these,…

Multimedia · Computer Science 2022-10-11 Chirag Raman , Jose Vargas-Quiros , Stephanie Tan , Ashraful Islam , Ekin Gedik , Hayley Hung