中文
相关论文

相关论文: ViDA-MAN: Visual Dialog with Digital Humans

200 篇论文

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Evonne Ng , Javier Romero , Timur Bagautdinov , Shaojie Bai , Trevor Darrell , Angjoo Kanazawa , Alexander Richard

We propose MAViD, a novel Multimodal framework for Audio-Visual Dialogue understanding and generation. Existing approaches primarily focus on non-interactive systems and are limited to producing constrained and unnatural human speech. The…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Youxin Pang , Jiajun Liu , Lingfeng Tan , Yong Zhang , Feng Gao , Xiang Deng , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Conversational AIs, or chatbots, mimic human speech when conversing. Smart assistants facilitate the automation of several tasks that needed human intervention earlier. Because of their accuracy, absence of dependence on human resources,…

计算与语言 · 计算机科学 2023-12-05 Aditya Paranjape , Yash Patwardhan , Vedant Deshpande , Aniket Darp , Jayashree Jagdale

Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privileged state or scripted control, which limits scalability to…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Hang Ye , Xiaoxuan Ma , Fan Lu , Wayne Wu , Kwan-Yee Lin , Yizhou Wang

We introduce VASA, a framework for generating lifelike talking faces with appealing visual affective skills (VAS) given a single static image and a speech audio clip. Our premiere model, VASA-1, is capable of not only generating lip…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sicheng Xu , Guojun Chen , Yu-Xiao Guo , Jiaolong Yang , Chong Li , Zhenyu Zang , Yizhong Zhang , Xin Tong , Baining Guo

With the arising concerns for the AI systems provided with direct access to abundant sensitive information, researchers seek to develop more reliable AI with implicit information sources. To this end, in this paper, we introduce a new task…

计算机视觉与模式识别 · 计算机科学 2020-08-25 Ye Zhu , Yu Wu , Yi Yang , Yan Yan

Human interactions are characterized by explicit as well as implicit channels of communication. While the explicit channel transmits overt messages, the implicit ones transmit hidden messages about the communicator (e.g., his/her intentions…

人机交互 · 计算机科学 2017-07-12 Katie Hoemann , Behnaz Rezaei , Stacy C. Marsella , Sarah Ostadabbas

We present a system that enables real-time interaction between human users and agents trained to control fighter jets in simulated 3D air combat scenarios. The agents are trained in a dedicated environment using Multi-Agent Reinforcement…

人工智能 · 计算机科学 2025-10-01 Ardian Selmonaj , Giacomo Del Rio , Adrian Schneider , Alessandro Antonucci

We present an end-to-end voice-based conversational agent that is able to engage in naturalistic multi-turn dialogue and align with the interlocutor's conversational style. The system uses a series of deep neural network components for…

人机交互 · 计算机科学 2019-08-15 Rens Hoegen , Deepali Aneja , Daniel McDuff , Mary Czerwinski

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond…

人工智能 · 计算机科学 2025-05-06 Yemin Shi , Yu Shu , Siwei Dong , Guangyi Liu , Jaward Sesay , Jingwen Li , Zhiting Hu

In the rapidly evolving landscape of human-robot collaboration, effective communication between humans and robots is crucial for complex task execution. Traditional request-response systems often lack naturalness and may hinder efficiency.…

机器人学 · 计算机科学 2024-09-12 Davide Ferrari , Filippo Alberi , Cristian Secchi

We present the design of an online social skills development interface for teenagers with autism spectrum disorder (ASD). The interface is intended to enable private conversation practice anywhere, anytime using a web-browser. Users…

The integration of large language models (LLMs) into virtual reality (VR) environments has opened new pathways for creating more immersive and interactive digital humans. By leveraging the generative capabilities of LLMs alongside…

This paper introduces a new model to generate rhythmically relevant non-verbal facial behaviors for virtual agents while they speak. The model demonstrates perceived performance comparable to behaviors directly extracted from the data and…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Alice Delbosc , Magalie Ochs , Nicolas Sabouret , Brian Ravenet , Stéphane Ayache

Holographic Intellectual Voice Assistant (HIVA) aims to facilitate human computer interaction using audiovisual effects and 3D avatar. HIVA provides complete information about the university, including requests of various nature: admission,…

人机交互 · 计算机科学 2023-07-13 Ruslan Isaev , Radmir Gumerov , Gulzada Esenalieva , Remudin Reshid Mekuria , Ermek Doszhanov

Modeling face-to-face communication in computer vision, which focuses on recognizing and analyzing nonverbal cues and behaviors during interactions, serves as the foundation for our proposed alternative to text-based Human-AI interaction.…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Dragos Costea , Alina Marcu , Cristina Lazar , Marius Leordeanu

Responsing with image has been recognized as an important capability for an intelligent conversational agent. Yet existing works only focus on exploring the multimodal dialogue models which depend on retrieval-based methods, but neglecting…

计算与语言 · 计算机科学 2022-03-30 Qingfeng Sun , Yujing Wang , Can Xu , Kai Zheng , Yaming Yang , Huang Hu , Fei Xu , Jessica Zhang , Xiubo Geng , Daxin Jiang

LLM-based translation agents have achieved highly human-like translation results and are capable of handling longer and more complex contexts with greater efficiency. However, they are typically limited to text-only inputs. In this paper,…

Virtual assistants (VAs) have become ubiquitous in daily life, integrated into smartphones and smart devices, sparking interest in AI companions that enhance user experiences and foster emotional connections. However, existing companions…

人机交互 · 计算机科学 2025-09-03 Xuetong Wang , Ching Christie Pang , Pan Hui

Visual dialog is a challenging vision-language task in which a series of questions visually grounded by a given image are answered. To resolve the visual dialog task, a high-level understanding of various multimodal inputs (e.g., question,…

人工智能 · 计算机科学 2020-10-08 Sungjin Park , Taesun Whang , Yeochan Yoon , Heuiseok Lim