中文
相关论文

相关论文: A High-Fidelity Open Embodied Avatar with Lip Sync…

200 篇论文

There is a growing demand for the accessible creation of high-quality 3D avatars that are animatable and customizable. Although 3D morphable models provide intuitive control for editing and animation, and robustness for single-view face…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Connor Z. Lin , Koki Nagano , Jan Kautz , Eric R. Chan , Umar Iqbal , Leonidas Guibas , Gordon Wetzstein , Sameh Khamis

The facial expression generation capability of humanoid social robots is critical for achieving natural and human-like interactions, playing a vital role in enhancing the fluidity of human-robot interactions and the accuracy of emotional…

机器人学 · 计算机科学 2025-10-28 Yongtong Zhu , Lei Li , Iggy Qian , WenBin Zhou , Ye Yuan , Qingdu Li , Na Liu , Jianwei Zhang

This paper introduces an open-source simulator, BeliefNest, designed to enable embodied agents to perform collaborative tasks by leveraging Theory of Mind. BeliefNest dynamically and hierarchically constructs simulators within a Minecraft…

人工智能 · 计算机科学 2025-05-20 Rikunari Sagara , Koichiro Terao , Naoto Iwahashi

Lip-to-Speech (Lip2Speech) synthesis, which predicts corresponding speech from talking face images, has witnessed significant progress with various models and training strategies in a series of independent studies. However, existing studies…

多媒体 · 计算机科学 2023-05-25 Zheng-Yan Sheng , Yang Ai , Zhen-Hua Ling

The emergence of commercial tools for real-time performance-based 2D animation has enabled 2D characters to appear on live broadcasts and streaming platforms. A key requirement for live animation is fast and accurate lip sync that allows…

图形学 · 计算机科学 2019-10-22 Deepali Aneja , Wilmot Li

This study explores a streamlined facial data collection method for conversational contexts, addressing the limitations of existing approaches that often require extensive datasets and prioritize technical metrics over user perception and…

人机交互 · 计算机科学 2026-02-03 Seoyoung Kang , Seokhwan Yang , Hail Song , Boram Yoon , Jinwook Kim , Kangsoo Kim , Woontack Woo

Embodied agents tasked with complex scenarios, whether in real or simulated environments, rely heavily on robust planning capabilities. When instructions are formulated in natural language, large language models (LLMs) equipped with…

The rapid advancement of Large Language Models (LLMs) has marked a significant breakthrough in Artificial Intelligence (AI), ushering in a new era of Human-centered Artificial Intelligence (HAI). HAI aims to better serve human welfare and…

机器人学 · 计算机科学 2025-10-29 Wenbin Ding , Jun Chen , Mingjia Chen , Fei Xie , Qi Mao , Philip Dames

This paper introduces the concept of coexistence for embodied artificial agents and argues that it is a prerequisite for long-term, in-the-wild interaction with humans. Contemporary embodied artificial agents excel in static, predefined…

机器学习 · 计算机科学 2025-06-03 Hannah Kuehn , Joseph La Delfa , Miguel Vasco , Danica Kragic , Iolanda Leite

Face-to-face conversation in Virtual Reality (VR) is a challenge when participants wear head-mounted displays (HMD). A significant portion of a participant's face is hidden and facial expressions are difficult to perceive. Past research has…

图形学 · 计算机科学 2020-11-10 Philipp Ladwig , Alexander Pech , Ralf Dörner , Christian Geiger

Drones operating in human-occupied spaces suffer from insufficient communication mechanisms that create uncertainty about their intentions. We present HoverAI, an embodied aerial agent that integrates drone mobility,…

Facial expression retargeting from humans to virtual characters is a useful technique in computer graphics and animation. Traditional methods use markers or blendshapes to construct a mapping between the human and avatar faces. However,…

计算机视觉与模式识别 · 计算机科学 2020-08-13 Juyong Zhang , Keyu Chen , Jianmin Zheng

Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelity lip-sync with different techniques, their task-orientated…

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Streaming speech-to-avatar synthesis creates real-time animations for a virtual character from audio data. Accurate avatar representations of speech are important for the visualization of sound in linguistics, phonetics, and phonology,…

声音 · 计算机科学 2023-10-26 Tejas S. Prabhune , Peter Wu , Bohan Yu , Gopala K. Anumanchipalli

SmartAvatar is a vision-language-agent-driven framework for generating fully rigged, animation-ready 3D human avatars from a single photo or textual prompt. While diffusion-based methods have made progress in general 3D object generation,…

计算机视觉与模式识别 · 计算机科学 2025-06-06 Alexander Huang-Menders , Xinhang Liu , Andy Xu , Yuyao Zhang , Chi-Keung Tang , Yu-Wing Tai

In order to create effective storytelling agents three fundamental questions must be answered: first, is a physically embodied agent preferable to a virtual agent or a voice-only narration? Second, does a human voice have an advantage over…

机器人学 · 计算机科学 2016-07-20 Sandra Costa , Alberto Brunete , Byung-Chull Bae , Nikolaos Mavridis

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Embodied artificial intelligence (Embodied AI) plays a pivotal role in the application of advanced technologies in the intelligent era, where AI systems are integrated with physical bodies that enable them to perceive, reason, and interact…

人工智能 · 计算机科学 2025-06-24 Zhaohan Feng , Ruiqi Xue , Lei Yuan , Yang Yu , Ning Ding , Meiqin Liu , Bingzhao Gao , Jian Sun , Xinhu Zheng , Gang Wang

Due to privacy concerns, open dialogue datasets for mental health are primarily generated through human or AI synthesis methods. However, the inherent implicit nature of psychological processes, particularly those of clients, poses…

人机交互 · 计算机科学 2025-11-18 Lixiu Wu , Yuanrong Tang , Qisen Pan , Xianyang Zhan , Yucheng Han , Lanxi Xiao , Tianhong Wang , Chen Zhong , Jiangtao Gong