English
Related papers

Related papers: Allo-AVA: A Large-Scale Multimodal Conversational …

200 papers

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

The global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multi-modal dataset for speech…

Sound · Computer Science 2025-07-08 Dongliang Zhou , Yakun Zhang , Jinghan Wu , Xingyu Zhang , Liang Xie , Erwei Yin

In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Wenqi Jia , Miao Liu , Hao Jiang , Ishwarya Ananthabhotla , James M. Rehg , Vamsi Krishna Ithapu , Ruohan Gao

We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms…

Graphics · Computer Science 2025-11-04 Jiye Lee , Chenghui Li , Linh Tran , Shih-En Wei , Jason Saragih , Alexander Richard , Hanbyul Joo , Shaojie Bai

Egocentric AI assistants in real-world settings must process multi-modal inputs (video, audio, text), respond in real time, and retain evolving long-term memory. However, existing benchmarks typically evaluate these abilities in isolation,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-11 Jiaqi Yan , Ruilong Ren , Jingren Liu , Shuning Xu , Ling Wang , Yiheng Wang , Xinlin Zhong , Yun Wang , Long Zhang , Xiangyu Chen , Changzhi Sun , Jixiang Luo , Dell Zhang , Hao Sun , Chi Zhang , Xuelong Li

LLaVA-Interactive is a research prototype for multimodal human-AI interaction. The system can have multi-turn dialogues with human users by taking multimodal user inputs and generating multimodal responses. Importantly, LLaVA-Interactive…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Wei-Ge Chen , Irina Spiridonova , Jianwei Yang , Jianfeng Gao , Chunyuan Li

The recent advancement of Vision Language Action (VLA) models has driven a critical demand for large scale egocentric datasets. However, existing datasets are often limited by short episode durations, typically spanning only a few minutes,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Senthil Palanisamy , Abhishek Anand , Satpal Singh Rathor , Pratyush Patnaik , Shubhanshu Khatana , Ekaksh Janweja

Autonomous driving, particularly navigating complex and unanticipated scenarios, demands sophisticated reasoning and planning capabilities. While Multi-modal Large Language Models (MLLMs) offer a promising avenue for this, their use has…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Hidehisa Arai , Keita Miwa , Kento Sasaki , Yu Yamaguchi , Kohei Watanabe , Shunsuke Aoki , Issei Yamamoto

Audio-visual automatic speech recognition (AV-ASR) is an extension of ASR that incorporates visual cues, often from the movements of a speaker's mouth. Unlike works that simply focus on the lip motion, we investigate the contribution of…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Valentin Gabeur , Paul Hongsuck Seo , Arsha Nagrani , Chen Sun , Karteek Alahari , Cordelia Schmid

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Youliang Zhang , Zhaoyang Li , Duomin Wang , Jiahe Zhang , Deyu Zhou , Zixin Yin , Xili Dai , Gang Yu , Xiu Li

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

Sound · Computer Science 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

We present the Multiview Extended Video with Activities (MEVA) dataset, a new and very-large-scale dataset for human activity recognition. Existing security datasets either focus on activity counts by aggregating public video disseminated…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Kellie Corona , Katie Osterdahl , Roderic Collins , Anthony Hoogs

Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communication is inherently…

Artificial Intelligence · Computer Science 2026-04-14 Yuzhe Weng , Haotian Wang , Xinyi Yu , Xiaoyan Wu , Haoran Xu , Shan He , Jun Du

VR Facial Animation is necessary in applications requiring clear view of the face, even though a VR headset is worn. In our case, we aim to animate the face of an operator who is controlling our robotic avatar system. We propose a real-time…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Andre Rochow , Max Schwarz , Michael Schreiber , Sven Behnke

Research interest in task-oriented dialogs has increased as systems such as Google Assistant, Alexa and Siri have become ubiquitous in everyday life. However, the impact of academic research in this area has been limited by the lack of…

Animating an avatar that reflects a user's action in the VR world enables natural interactions with the virtual environment. It has the potential to allow remote users to communicate and collaborate in a way as if they met in person.…

Graphics · Computer Science 2022-09-14 Yongjing Ye , Libin Liu , Lei Hu , Shihong Xia

Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framework that uses both audio and visual cues to track the target…

Computation and Language · Computer Science 2025-11-17 Tuochao Chen , Bandhav Veluri , Hongyu Gong , Shyamnath Gollakota

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Mengyi Shan , Shouchieh Chang , Ziqian Bai , Shichen Liu , Yinda Zhang , Luchuan Song , Rohit Pandey , Sean Fanello , Zeng Huang

The domain of 3D talking head generation has witnessed significant progress in recent years. A notable challenge in this field consists in blending speech-related motions with expression dynamics, which is primarily caused by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Federico Nocentini , Claudio Ferrari , Stefano Berretti

Creating realistic, fully animatable whole-body avatars from a single portrait is challenging due to limitations in capturing subtle expressions, body movements, and dynamic backgrounds. Current evaluation datasets and metrics fall short in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Chaoyi Wang , Yifan Yang , Jun Pei , Lijie Xia , Jianpo Liu , Xiaobing Yuan , Xinhan Di