English
Related papers

Related papers: ARIG: Autoregressive Interactive Head Generation f…

200 papers

Human behaviors in real-world environments are inherently interactive, with an individual's motion shaped by surrounding agents and the scene. Such capabilities are essential for applications in virtual avatars, interactive animation, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Yaoqin Ye , Yiteng Xu , Qin Sun , Xinge Zhu , Yujing Sun , Yuexin Ma

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Lele Chen , Guofeng Cui , Celong Liu , Zhong Li , Ziyi Kou , Yi Xu , Chenliang Xu

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

Recently, autoregressive (AR) language models have emerged as a dominant approach in speech synthesis, offering expressive generation and scalable training. However, conventional AR speech synthesis models relying on the next-token…

Sound · Computer Science 2025-06-30 Bohan Li , Zhihan Li , Haoran Wang , Hanglei Zhang , Yiwei Guo , Hankun Wang , Xie Chen , Kai Yu

Human dialogue contains evolving concepts, and speakers naturally associate multiple concepts to compose a response. However, current dialogue models with the seq2seq framework lack the ability to effectively manage concept transitions and…

Computation and Language · Computer Science 2021-09-10 Yicheng Zou , Zhihua Liu , Xingwu Hu , Qi Zhang

Real-world perception and interaction are inherently multimodal, encompassing not only language but also vision and speech, which motivates the development of "Omni" MLLMs that support both multimodal inputs and multimodal outputs. While a…

Machine Learning · Computer Science 2026-01-27 Dongjie Cheng , Ruifeng Yuan , Yongqi Li , Runyang You , Wenjie Wang , Liqiang Nie , Lei Zhang , Wenjie Li

Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kesen Zhao , Jiaxin Shi , Beier Zhu , Junbao Zhou , Xiaolong Shen , Yuan Zhou , Qianru Sun , Hanwang Zhang

Conventional wisdom suggests that autoregressive models are used to process discrete data. When applied to continuous modalities such as visual data, Visual AutoRegressive modeling (VAR) typically resorts to quantization-based approaches to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chenze Shao , Fandong Meng , Jie Zhou

Retrieval-augmented generation (RAG) techniques have proven to be effective in integrating up-to-date information, mitigating hallucinations, and enhancing response quality, particularly in specialized domains. While many RAG approaches…

Retrieval-augmented generation (RAG) systems combine the strengths of language generation and information retrieval to power many real-world applications like chatbots. Use of RAG for understanding of videos is appealing but there are two…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Md Adnan Arefeen , Biplob Debnath , Md Yusuf Sarwar Uddin , Srimat Chakradhar

Articulation, emotion, and personality play strong roles in the orofacial movements. To improve the naturalness and expressiveness of virtual agents (VAs), it is important that we carefully model the complex interplay between these factors.…

Human-Computer Interaction · Computer Science 2023-05-15 Najmeh Sadoughi , Carlos Busso

We propose a real-time system for synthesizing gestures directly from speech. Our data-driven approach is based on Generative Adversarial Neural Networks to model the speech-gesture relationship. We utilize the large amount of speaker video…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Manuel Rebol , Christian Gütl , Krzysztof Pietroszek

Time series modeling is crucial for many applications, however, it faces challenges such as complex spatio-temporal dependencies and distribution shifts in learning from historical context to predict task-specific outcomes. To address these…

Artificial Intelligence · Computer Science 2024-08-28 Chidaksh Ravuru , Sagar Srinivas Sakhinana , Venkataramana Runkana

Flexible and natural nonverbal reactions to human behavior remain a challenge for socially interactive agents (SIAs) that are predominantly animated using hand-crafted rules. While recently proposed machine learning based approaches to…

Human-Computer Interaction · Computer Science 2024-02-14 Daksitha Withanage Don , Philipp Müller , Fabrizio Nunnari , Elisabeth André , Patrick Gebhard

Conversational scenarios are very common in real-world settings, yet existing co-speech motion synthesis approaches often fall short in these contexts, where one person's audio and gestures will influence the other's responses.…

Sound · Computer Science 2024-12-04 Mingyi Shi , Dafei Qin , Leo Ho , Zhouyingcheng Liao , Yinghao Huang , Junichi Yamagishi , Taku Komura

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical…

Human-Computer Interaction · Computer Science 2025-08-05 Kazushi Kato , Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

Flow models are effective at progressively generating realistic images, but they generally struggle to capture long-range dependencies during the generation process as they compress all the information from previous time steps into a single…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Mude Hui , Rui-Jie Zhu , Songlin Yang , Yu Zhang , Zirui Wang , Yuyin Zhou , Jason Eshraghian , Cihang Xie

The advent of artificial intelligence has contributed in a groundbreaking transformation of the fashion industry, redefining creativity and innovation in unprecedented ways. This work investigates methodologies for generating tailored…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Georgia Argyrou , Angeliki Dimitriou , Maria Lymperaiou , Giorgos Filandrianos , Giorgos Stamou

Egocentric interactive world models are essential for augmented reality and embodied AI, where visual generation must respond to user input with low latency, geometric consistency, and long-term stability. We study egocentric interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Yuxi Wang , Wenqi Ouyang , Tianyi Wei , Yi Dong , Zhiqi Shen , Xingang Pan
‹ Prev 1 3 4 5 6 7 10 Next ›