中文
相关论文

相关论文: X-Streamer: Unified Human World Modeling with Audi…

200 篇论文

Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Sreehari Rajan , Kunal Bhosikar , Charu Sharma

We present X-Avatar, a novel avatar model that captures the full expressiveness of digital humans to bring about life-like experiences in telepresence, AR/VR and beyond. Our method models bodies, hands, facial expressions and appearance in…

计算机视觉与模式识别 · 计算机科学 2023-03-10 Kaiyue Shen , Chen Guo , Manuel Kaufmann , Juan Jose Zarate , Julien Valentin , Jie Song , Otmar Hilliges

Video reasoning constitutes a comprehensive assessment of a model's capabilities, as it demands robust perceptual and interpretive skills, thereby serving as a means to explore the boundaries of model performance. While recent research has…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Yudi Shi , Shangzhe Di , Qirui Chen , Qinian Wang , Jiayin Cai , Xiaolong Jiang , Yao Hu , Weidi Xie

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that generates realistic, context-aware dynamics for both the…

Existing video avatar models can produce fluid human animations, yet they struggle to move beyond mere physical likeness to capture a character's authentic essence. Their motions typically synchronize with low-level cues like audio rhythm,…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Jianwen Jiang , Weihong Zeng , Zerong Zheng , Jiaqi Yang , Chao Liang , Wang Liao , Han Liang , Yuan Zhang , Mingyuan Gao

The landscape of video generation is shifting, from a focus on generating visually appealing clips to building virtual environments that support interaction and maintain physical plausibility. These developments point toward the emergence…

人工智能 · 计算机科学 2026-02-09 Jingtong Yue , Ziqi Huang , Zhaoxi Chen , Xintao Wang , Pengfei Wan , Ziwei Liu

The problem of building a coherent and non-monotonous conversational agent with proper discourse and coverage is still an area of open research. Current architectures only take care of semantic and contextual information for a given query…

计算与语言 · 计算机科学 2025-04-22 Gaurav Kumar , Rishabh Joshi , Jaspreet Singh , Promod Yenigalla

Existing end-to-end sign-language animation systems suffer from low naturalness, limited facial/body expressivity, and no user control. We propose a human-centered, real-time speech-to-sign animation framework that integrates (1) a…

人机交互 · 计算机科学 2025-06-25 Yingchao Li

Effectively handling the interplay between spatial perception and action generation remains a critical bottleneck in robotic manipulation. Existing methods typically treat spatial perception and action execution as decoupled or strictly…

机器人学 · 计算机科学 2026-05-13 Kai Xiong , Hongjie Fang , Lixin Yang , Cewu Lu

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

In-person human interaction relies on our spatial perception of each other and our surroundings. Current remote communication tools partially address each of these aspects. Video calls convey real user representations but without spatial…

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that…

计算机视觉与模式识别 · 计算机科学 2023-05-19 Lihui Qian , Xintong Han , Faqiang Wang , Hongyu Liu , Haoye Dong , Zhiwen Li , Huawei Wei , Zhe Lin , Cheng-Bin Jin

Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates…

人工智能 · 计算机科学 2024-11-06 Zhifei Xie , Changqiao Wu

While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coherence, which leads to…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Yuxuan Luo , Zhengkun Rong , Lizhen Wang , Longhao Zhang , Tianshu Hu , Yongming Zhu

In this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Jiale Zhang , Yulun Zhang , Jinjin Gu , Jiahua Dong , Linghe Kong , Xiaokang Yang

Recent diffusion-based human image animation techniques have demonstrated impressive success in synthesizing videos that faithfully follow a given reference identity and a sequence of desired movement poses. Despite this, there are still…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Xiang Wang , Shiwei Zhang , Changxin Gao , Jiayu Wang , Xiaoqiang Zhou , Yingya Zhang , Luxin Yan , Nong Sang

In this paper we present a Transformer-Transducer model architecture and a training technique to unify streaming and non-streaming speech recognition models into one model. The model is composed of a stack of transformer layers for audio…

声音 · 计算机科学 2020-10-08 Anshuman Tripathi , Jaeyoung Kim , Qian Zhang , Han Lu , Hasim Sak

Realistic and interactive traffic simulation is essential for training and evaluating autonomous driving systems. However, most existing data-driven simulation methods rely on static initialization or log-replay data, limiting their ability…

机器人学 · 计算机科学 2026-03-04 Zhenghao Peng , Yuxin Liu , Bolei Zhou

Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Learning such performance from video is a promising alternative to traditional 3D pipelines.…

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot architecture that enables…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yawen Luo , Xiaoyu Shi , Junhao Zhuang , Yutian Chen , Quande Liu , Xintao Wang , Pengfei Wan , Tianfan Xue