中文
相关论文

相关论文: X-Streamer: Unified Human World Modeling with Audi…

200 篇论文

We present X-Actor, a novel audio-driven portrait animation framework that generates lifelike, emotionally expressive talking head videos from a single reference image and an input audio clip. Unlike prior methods that emphasize lip…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Chenxu Zhang , Zenan Li , Hongyi Xu , You Xie , Xiaochen Zhao , Tianpei Gu , Guoxian Song , Xin Chen , Chao Liang , Jianwen Jiang , Linjie Luo

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zhiyao Sun , Ziqiao Peng , Yifeng Ma , Yi Chen , Zhengguang Zhou , Zixiang Zhou , Guozhen Zhang , Youliang Zhang , Yuan Zhou , Qinglin Lu , Yong-Jin Liu

We present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-diffusion framework,…

计算机视觉与模式识别 · 计算机科学 2025-07-14 Zeyuan Chen , Hongyi Xu , Guoxian Song , You Xie , Chenxu Zhang , Xin Chen , Chao Wang , Di Chang , Linjie Luo

In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming…

Existing video avatar models have demonstrated impressive capabilities in scenarios such as talking, public speaking, and singing. However, the majority of these methods exhibit limited alignment with respect to text instructions,…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Ruikui Wang , Jinheng Feng , Lang Tian , Huaishao Luo , Chaochao Li , Liangbo Zhou , Huan Zhang , Youzheng Wu , Xiaodong He

Understanding continuous video streams plays a fundamental role in real-time applications including embodied AI and autonomous driving. Unlike offline video understanding, streaming video understanding requires the ability to process video…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Yibin Yan , Jilan Xu , Shangzhe Di , Yikun Liu , Yudi Shi , Qirui Chen , Zeqian Li , Yifei Huang , Weidi Xie

Creating AI systems that can interact with environments over long periods, similar to human cognition, has been a longstanding research goal. Recent advancements in multimodal large language models (MLLMs) have made significant strides in…

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Enric Corona , Andrei Zanfir , Eduard Gabriel Bazavan , Nikos Kolotouros , Thiemo Alldieck , Cristian Sminchisescu

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

Modern visual agents require representations that are general, causal, and physically structured to operate in real-time streaming environments. However, current vision foundation models remain fragmented, specializing narrowly in image…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Yibin Yan , Jilan Xu , Shangzhe Di , Haoning Wu , Weidi Xie

Real-time understanding of continuous video streams is essential for interactive assistants and multimodal agents operating in dynamic environments. However, most existing video reasoning approaches follow a batch paradigm that defers…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Zikang Liu , Longteng Guo , Handong Li , Ru Zhen , Xingjian He , Ruyi Ji , Xiaoming Ren , Yanhao Zhang , Haonan Lu , Jing Liu

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yu Zhang , Kaiyuan Shen , Yang Li

Recent advancements in neural rendering technologies and their supporting devices have paved the way for immersive 3D experiences, significantly transforming human interaction with intelligent devices across diverse applications. However,…

图形学 · 计算机科学 2025-04-01 Chaojian Li , Sixu Li , Linrui Jiang , Jingqun Zhang , Yingyan Celine Lin

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Xiang Deng , Feng Gao , Yong Zhang , Youxin Pang , Xu Xiaoming , Zhuoliang Kang , Xiaoming Wei , Yebin Liu

Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing this perceptual pathway is critical for building natural full-duplex interactive systems. We propose…

We present an architecture for integrating real-time, multimodal input into a computational agent's contextual model. Using a human-avatar interaction in a virtual world, we treat aligned gesture and speech as an ensemble where content may…

人机交互 · 计算机科学 2019-09-19 Nikhil Krishnaswamy , James Pustejovsky

Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams to driving actions. Yet, current evaluation pipelines still…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Chaoda Zheng , Sean Li , Jinhao Deng , Zhennan Wang , Shijia Chen , Liqiang Xiao , Ziheng Chi , Hongbin Lin , Kangjie Chen , Boyang Wang , Yu Zhang , Xianming Liu

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Gaojie Lin , Jianwen Jiang , Jiaqi Yang , Zerong Zheng , Chao Liang

Benefiting from the advancements in large language models and cross-modal alignment, existing multi-modal video understanding methods have achieved prominent performance in offline scenario. However, online video streams, as one of the most…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Haoji Zhang , Yiqin Wang , Yansong Tang , Yong Liu , Jiashi Feng , Jifeng Dai , Xiaojie Jin

Active Real-time interaction with video LLMs introduces a new paradigm for human-computer interaction, where the model not only understands user intent but also responds while continuously processing streaming video on the fly. Unlike…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Rui Qian , Shuangrui Ding , Xiaoyi Dong , Pan Zhang , Yuhang Zang , Yuhang Cao , Dahua Lin , Jiaqi Wang
‹ 上一页 1 2 3 10 下一页 ›