中文
相关论文

相关论文: RITA: A Real-time Interactive Talking Avatars Fram…

200 篇论文

Creativity in AI imagery remains a fundamental challenge, requiring not only the generation of visually compelling content but also the capacity to add novel, expressive, and artistically rich transformations to images. Unlike conventional…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Kavana Venkatesh , Connor Dunlop , Pinar Yanardag

A widespread adoption of Virtual, Augmented, and Mixed Reality (VR/AR/MR), collectively referred to as Extended Reality (XR), has become a tangible possibility to revolutionize educational and training scenarios by offering immersive,…

We introduce TADA, a simple-yet-effective approach that takes textual descriptions and produces expressive 3D avatars with high-quality geometry and lifelike textures, that can be animated and rendered with traditional graphics pipelines.…

人工智能 · 计算机科学 2023-08-22 Tingting Liao , Hongwei Yi , Yuliang Xiu , Jiaxaing Tang , Yangyi Huang , Justus Thies , Michael J. Black

The Metaverse represents a transformative shift beyond traditional mobile Internet, creating an immersive, persistent digital ecosystem where users can interact, socialize, and work within 3D virtual environments. Powered by large models…

计算机与社会 · 计算机科学 2025-08-12 Yuntao Wang , Qinnan Hu , Zhou Su , Linkang Du , Qichao Xu , Weiwei Li

We present a novel approach for generating plausible verbal interactions between virtual human-like agents and user avatars in shared virtual environments. Sense-Plan-Ask, or SPA, extends prior work in propositional planning and natural…

多智能体系统 · 计算机科学 2020-02-11 Andrew Best , Sahil Narang , Dinesh Manocha

An engaging and provocative question can open up a great conversation. In this work, we explore a novel scenario: a conversation agent views a set of the user's photos (for example, from social media platforms) and asks an engaging question…

人工智能 · 计算机科学 2022-05-20 Shih-Han Chan , Tsai-Lun Yang , Yun-Wei Chu , Chi-Yang Hsu , Ting-Hao Huang , Yu-Shian Chiu , Lun-Wei Ku

Instruction-guided image editing offers an intuitive way for users to edit images with natural language. However, diffusion-based editing models often struggle to accurately interpret complex user instructions, especially those involving…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Ziyun Zeng , Hang Hua , Jiebo Luo

This paper presents Matrix, an advanced AI-powered framework designed for real-time 3D object generation in Augmented Reality (AR) environments. By integrating a cutting-edge text-to-3D generative AI model, multilingual speech-to-text…

人机交互 · 计算机科学 2025-03-24 Majid Behravan , Denis Gracanin

Traditional animation production involves complex pipelines and significant manual labor cost. While recent video generation models such as Sora, Kling, and CogVideoX achieve impressive results on natural video synthesis, they exhibit…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Ruiqi Chen , Kaitong Cai , Yijia Fan , Keze Wang

Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relations between audio events described in the input text. However,…

机器学习 · 计算机科学 2026-04-10 Yuhang He , Yash Jain , Xubo Liu , Andrew Markham , Vibhav Vineet

In recent years, facial video generation models have gained popularity. However, these models often lack expressive power when dealing with exaggerated anime-style faces due to the absence of high-quality anime-style face training sets. We…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Xiaokai Chen , Xuan Liu , Donglin Di , Yongjia Ma , Wei Chen , Tonghua Su

In this study, our goal is to create interactive avatar agents that can autonomously plan and animate nuanced facial movements realistically, from both visual and behavioral perspectives. Given high-level inputs about the environment and…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Duomin Wang , Bin Dai , Yu Deng , Baoyuan Wang

Proactive human-robot interaction (HRI) allows the receptionist robots to actively greet people and offer services based on vision, which has been found to improve acceptability and customer satisfaction. Existing approaches are either…

机器人学 · 计算机科学 2021-03-25 Yang Xue , Fan Wang , Hao Tian , Min Zhao , Jiangyong Li , Haiqing Pan , Yueqiang Dong

We present RealityTalk, a system that augments real-time live presentations with speech-driven interactive virtual elements. Augmented presentations leverage embedded visuals and animation for engaging and expressive storytelling. However,…

人机交互 · 计算机科学 2022-08-15 Jian Liao , Adnan Karim , Shivesh Jadon , Rubaiat Habib Kazi , Ryo Suzuki

Face Recognition (FR) has advanced significantly with the development of deep learning, achieving high accuracy in several applications. However, the lack of interpretability of these systems raises concerns about their accountability,…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Ivan DeAndres-Tame , Muhammad Faisal , Ruben Tolosana , Rouqaiah Al-Refai , Ruben Vera-Rodriguez , Philipp Terhörst

Social chatbots have gained immense popularity, and their appeal lies not just in their capacity to respond to the diverse requests from users, but also in the ability to develop an emotional connection with users. To further develop and…

计算与语言 · 计算机科学 2022-06-01 Avinash Madasu , Mauajama Firdaus , Asif Eqbal

Drawing is an art that enables people to express their imagination and emotions. However, individuals usually face challenges in drawing, especially when translating conceptual ideas into visually coherent representations and bridging the…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Trong-Vu Hoang , Quang-Binh Nguyen , Duy-Nam Ly , Khanh-Duy Le , Tam V. Nguyen , Minh-Triet Tran , Trung-Nghia Le

Conversational systems must be robust to user interactions that naturally exhibit diverse conversational traits. Capturing and simulating these diverse traits coherently and efficiently presents a complex challenge. This paper introduces…

计算与语言 · 计算机科学 2024-10-29 Rafael Ferreira , David Semedo , João Magalhães

To improve the experiences of face-to-face conversation with avatar, this paper presents a novel conversation system. It is composed of two sequence-to-sequence models respectively for listening and speaking and a Generative Adversarial…

计算机视觉与模式识别 · 计算机科学 2019-08-22 Zezhou Chen , Zhaoxiang Liu , Huan Hu , Jinqiang Bai , Shiguo Lian , Fuyuan Shi , Kai Wang

Modeling face-to-face communication in computer vision, which focuses on recognizing and analyzing nonverbal cues and behaviors during interactions, serves as the foundation for our proposed alternative to text-based Human-AI interaction.…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Dragos Costea , Alina Marcu , Cristina Lazar , Marius Leordeanu