中文
相关论文

相关论文: X-Streamer: Unified Human World Modeling with Audi…

200 篇论文

Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communication is inherently…

人工智能 · 计算机科学 2026-04-14 Yuzhe Weng , Haotian Wang , Xinyi Yu , Xiaoyan Wu , Haoran Xu , Shan He , Jun Du

Multimodal large language models (MLLMs) have shown strong performance on offline video understanding, but most are limited to offline inference or have weak online reasoning, making multi-turn interaction over continuously arriving video…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Lu Wang , Zhuoran Jin , Yupu Hao , Yubo Chen , Kang Liu , Yulong Ao , Jun Zhao

Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Yuxiao Yang , Hualian Sheng , Sijia Cai , Jing Lin , Jiahao Wang , Bing Deng , Junzhe Lu , Haoqian Wang , Jieping Ye

Speech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to…

计算机视觉与模式识别 · 计算机科学 2024-02-09 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Chen Qian , Zhaoxiang Zhang , Zhen Lei

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn…

With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited exploration of unified generative architectures. In addition, many…

多媒体 · 计算机科学 2025-07-08 Lei Zhao , Linfeng Feng , Dongxu Ge , Rujin Chen , Fangqiu Yi , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

Video generation is an inherently challenging task, as it requires modeling realistic temporal dynamics as well as spatial content. Existing methods entangle the two intrinsically different tasks of motion and content creation in a single…

计算机视觉与模式识别 · 计算机科学 2020-01-13 Ximeng Sun , Huijuan Xu , Kate Saenko

Extended reality (XR) is rapidly advancing, and poised to revolutionize content creation and consumption. In XR, users integrate various sensory inputs to form a cohesive perception of the virtual environment. This survey reviews the…

多媒体 · 计算机科学 2025-06-13 Haopeng Wang , Haiwei Dong , Abdulmotaleb El Saddik

Realistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather…

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Aggelina Chatziagapi , Louis-Philippe Morency , Hongyu Gong , Michael Zollhoefer , Dimitris Samaras , Alexander Richard

Lane segment topology reasoning constructs a comprehensive road network by capturing the topological relationships between lane segments and their semantic types. This enables end-to-end autonomous driving systems to perform road-dependent…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Yiming Yang , Yueru Luo , Bingkun He , Hongbin Lin , Suzhong Fu , Chao Zheng , Zhipeng Cao , Erlong Li , Chao Yan , Shuguang Cui , Zhen Li

Swim extends the actor model to support applications composed of linked distributed actors that continuously analyze boundless streams of events from millions of sources, to respond in-sync with the real-world. Swim builds a running…

分布式、并行与集群计算 · 计算机科学 2022-05-24 Chris Sachs , Ajay Govindarajan , Simon Crosby

Diffusion models have shown impressive potential on talking head generation. While plausible appearance and talking effect are achieved, these methods still suffer from temporal, 3D or expression inconsistency due to the error accumulation…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Haijie Yang , Zhenyu Zhang , Hao Tang , Jianjun Qian , Jian Yang

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatenate representation of…

人工智能 · 计算机科学 2025-06-24 Shaolei Zhang , Shoutao Guo , Qingkai Fang , Yan Zhou , Yang Feng

We present a target-aware video diffusion model that generates videos from an input image, in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask, and the action…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Taeksoo Kim , Hanbyul Joo

In this paper, we present the design of a multimodal interaction framework for intelligent virtual agents in wearable mixed reality environments, especially for interactive applications at museums, botanical gardens, and similar places.…

人机交互 · 计算机科学 2025-03-26 Ghazanfar Ali , Hong-Quan Le , Junho Kim , Seoung-won Hwang , Jae-In Hwang

In e-commerce and digital marketing, generating high-fidelity human-product demonstration videos is important for effective product presentation. However, most existing frameworks either fail to preserve the identities of both humans and…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Lizhen Wang , Zhurong Xia , Tianshu Hu , Pengrui Wang , Pengfei Wei , Zerong Zheng , Ming Zhou , Yuan Zhang , Mingyuan Gao

Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools. However, existing works focus on (1) image-sharing behavior in singular sessions, leading to limited long-term…

计算与语言 · 计算机科学 2024-07-08 Young-Jun Lee , Dokyong Lee , Junyoung Youn , Kyeongjin Oh , Byungsoo Ko , Jonghwan Hyeon , Ho-Jin Choi

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and free-viewpoint talking portrait synthesis, given an identity embedding or reference image,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Foivos Paraperas Papantoniou , Stathis Galanakis , Rolandos Alexandros Potamias , Bernhard Kainz , Stefanos Zafeiriou

In recent years, audio-driven 3D facial animation has gained significant attention, particularly in applications such as virtual reality, gaming, and video conferencing. However, accurately modeling the intricate and subtle dynamics of…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Guinan Su , Yanwu Yang , Zhifeng Li