中文
相关论文

相关论文: ViDA-MAN: Visual Dialog with Digital Humans

200 篇论文

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

图形学 · 计算机科学 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

多媒体 · 计算机科学 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Most mobile health apps employ data visualization to help people view their health and activity data, but these apps provide limited support for visual data exploration. Furthermore, despite its huge potential benefits, mobile visualization…

人机交互 · 计算机科学 2021-01-19 Young-Ho Kim , Bongshin Lee , Arjun Srinivasan , Eun Kyoung Choe

Robots navigating in human environments should use language to ask for assistance and be able to understand human responses. To study this challenge, we introduce Cooperative Vision-and-Dialog Navigation, a dataset of over 2k embodied,…

计算与语言 · 计算机科学 2019-10-15 Jesse Thomason , Michael Murray , Maya Cakmak , Luke Zettlemoyer

Existing for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head movements and the structural complexity of hands. To address…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Donglin Huang , Yongyuan Li , Tianhang Liu , Junming Huang , Xiaoda Yang , Chi Wang , Weiwei Xu

AI enabled chat bots have recently been put to use to answer customer service queries, however it is a common feedback of users that bots lack a personal touch and are often unable to understand the real intent of the user's question. To…

This software project based paper is for a vision of the near future in which computer interaction is characterized by natural face-to-face conversations with lifelike characters that speak, emote, and gesture. The first step is speech. The…

人机交互 · 计算机科学 2013-05-10 Urmila Shrawankar , Anjali Mahajan

In this study, we propose a solution based on a multi-agent LLM architecture and a voice user interface (VUI) designed to update the knowledge base of a digital assistant. Its usability is evaluated in comparison to a more traditional…

人机交互 · 计算机科学 2025-05-29 Grzegorz Wolny , Michał Szczerbak

Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a…

计算机视觉与模式识别 · 计算机科学 2021-03-04 Shijie Geng , Peng Gao , Moitreya Chatterjee , Chiori Hori , Jonathan Le Roux , Yongfeng Zhang , Hongsheng Li , Anoop Cherian

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Visual Language Navigation is a task that challenges robots to navigate in realistic environments based on natural language instructions. While previous research has largely focused on static settings, real-world navigation must often…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Dillon Loh , Tomasz Bednarz , Xinxing Xia , Frank Guan

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Lavisha Aggarwal , Vikas Bahirwani , Andrea Colaco

Humans and other intelligent animals evolved highly sophisticated perception systems that combine multiple sensory modalities. On the other hand, state-of-the-art artificial agents rely mostly on visual inputs or structured low-dimensional…

机器学习 · 计算机科学 2021-07-07 Shashank Hegde , Anssi Kanervisto , Aleksei Petrenko

Intelligent agents have great potential as facilitators of group conversation among older adults. However, little is known about how to design agents for this purpose and user group, especially in terms of agent embodiment. To this end, we…

机器人学 · 计算机科学 2022-12-09 Katie Seaborn , Takuya Sekiguchi , Seiki Tokunaga , Norihisa P. Miyake , Mihoko Otake-Matsuura

Offering diverse perspectives on a museum artifact can deepen visitors' understanding and help avoid the cognitive limitations of a single narrative, ultimately enhancing their overall experience. Physical museums promote diversity through…

人机交互 · 计算机科学 2025-08-12 Mingyang Su , Chao Liu , Jingling Zhang , WU Shuang , Mingming Fan

With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience…

Speech recognition is the technology that enables machines to interpret and process human speech, converting spoken language into text or commands. This technology is essential for applications such as virtual assistants, transcription…

音频与语音处理 · 电气工程与系统科学 2025-01-09 Xinyu Wang , Haotian Jiang , Haolin Huang , Yu Fang , Mengjie Xu , Qian Wang

Dialogue systems have the potential to change how people interact with machines but are highly dependent on the quality of the data used to train them. It is therefore important to develop good dialogue annotation tools which can improve…

计算与语言 · 计算机科学 2019-11-06 Edward Collins , Nikolai Rozanov , Bingbing Zhang

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zeyu Zhu , Weijia Wu , Mike Zheng Shou