中文
相关论文

相关论文: Unified Multi-Modal Interactive & Reactive 3D Moti…

200 篇论文

Previous dominant methods for scene flow estimation focus mainly on input from two consecutive frames, neglecting valuable information in the temporal domain. While recent trends shift towards multi-frame reasoning, they suffer from rapidly…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Qingwen Zhang , Xiaomeng Zhu , Yushan Zhang , Yixi Cai , Olov Andersson , Patric Jensfelt

Humans perform a variety of interactive motions, among which duet dance is one of the most challenging interactions. However, in terms of human motion generative models, existing works are still unable to generate high-quality interactive…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Ronghui Li , Youliang Zhang , Yachao Zhang , Yuxiang Zhang , Mingyang Su , Jie Guo , Ziwei Liu , Yebin Liu , Xiu Li

Joint audio-video (AV) generation is still a significant challenge in generative AI, primarily due to three critical requirements: quality of the generated samples, seamless multimodal synchronization and temporal coherence, with audio…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Alex Ergasti , Giuseppe Gabriele Tarollo , Filippo Botti , Tomaso Fontanini , Claudio Ferrari , Massimo Bertozzi , Andrea Prati

Unbounded 3D world generation is emerging as a foundational task for scene modeling in computer vision, graphics, and robotics. In this work, we present WorldFlow3D, a novel method capable of generating unbounded 3D worlds. Building upon a…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Amogh Joshi , Julian Ost , Felix Heide

Generating realistic human motion with high-level controls is a crucial task for social understanding, robotics, and animation. With high-quality MOCAP data becoming more available recently, a wide range of data-driven approaches have been…

图形学 · 计算机科学 2025-07-29 Wenning Xu , Shiyu Fan , Paul Henderson , Edmond S. L. Ho

To effectively engage in human society, the ability to adapt, filter information, and make informed decisions in ever-changing situations is critical. As robots and intelligent agents become more integrated into human life, there is a…

Accurate 3D scene flow estimation is critical for autonomous systems to navigate dynamic environments safely, but creating the necessary large-scale, manually annotated datasets remains a significant bottleneck for developing robust…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Ajinkya Khoche , Qingwen Zhang , Yixi Cai , Sina Sharif Mansouri , Patric Jensfelt

This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Muhammad Islam , Tao Huang , Euijoon Ahn , Usman Naseem

Recently, the rectified flow (RF) has emerged as the new state-of-the-art among flow-based diffusion models due to its high efficiency advantage in straight path sampling, especially with the amazing images generated by a series of RF…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Zhiyuan Ma , Ruixun Liu , Sixian Liu , Jianjun Li , Bowen Zhou

Continual learning in robotics seeks systems that can constantly adapt to changing environments and tasks, mirroring human adaptability. A key challenge is refining dynamics models, essential for planning and control, while addressing…

机器人学 · 计算机科学 2025-09-09 Alejandro Murillo-Gonzalez , Lantao Liu

Realistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to basic material types…

计算机视觉与模式识别 · 计算机科学 2025-05-09 Zhuoman Liu , Weicai Ye , Yan Luximon , Pengfei Wan , Di Zhang

Text-to-Motion generation has become a fundamental task in human-machine interaction, enabling the synthesis of realistic human motions from natural language descriptions. Although recent advances in large language models and reinforcement…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Runqi Ouyang , Haoyun Li , Zhenyuan Zhang , Xiaofeng Wang , Zeyu Zhang , Zheng Zhu , Guan Huang , Sirui Han , Xingang Wang

Audio-driven bimanual piano motion generation requires precise modeling of complex musical structures and dynamic cross-hand coordination. However, existing methods often rely on acoustic-only representations lacking symbolic priors, employ…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Xuan Wang , Kai Ruan , Jiayi Han , Kaiyue Zhou , Gaoang Wang

Existing dominant methods for audio generation include Generative Adversarial Networks (GANs) and diffusion-based methods like Flow Matching. GANs suffer from slow convergence during training, while diffusion methods require multi-step…

音频与语音处理 · 电气工程与系统科学 2026-03-10 Zengwei Yao , Wei Kang , Han Zhu , Liyong Guo , Lingxuan Ye , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Long Lin , Daniel Povey

We propose GoalFlow, an end-to-end autonomous driving method for generating high-quality multimodal trajectories. In autonomous driving scenarios, there is rarely a single suitable trajectory. Recent methods have increasingly focused on…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zebin Xing , Xingyu Zhang , Yang Hu , Bo Jiang , Tong He , Qian Zhang , Xiaoxiao Long , Wei Yin

Large language models perform well in short text generation but still struggle with long text generation, particularly under complex constraints. Such tasks involve multiple tightly coupled objectives, including global structural…

计算与语言 · 计算机科学 2026-03-06 Yifan Zhu , Guanting Chen , Bing Wei , Haoran Luo

Synthesizing realistic human-object interaction motions is a critical problem in VR/AR and human animation. Unlike the commonly studied scenarios involving a single human or hand interacting with one object, we address a more generic…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Wenkun He , Yun Liu , Ruitao Liu , Li Yi

Obtaining the ground truth labels from a video is challenging since the manual annotation of pixel-wise flow labels is prohibitively expensive and laborious. Besides, existing approaches try to adapt the trained model on synthetic datasets…

计算机视觉与模式识别 · 计算机科学 2022-07-25 Yunhui Han , Kunming Luo , Ao Luo , Jiangyu Liu , Haoqiang Fan , Guiming Luo , Shuaicheng Liu

We present a novel method to generate human motion to populate 3D indoor scenes. It can be controlled with various combinations of conditioning signals such as a path in a scene, target poses, past motions, and scenes represented as 3D…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Nicolas Ugrinovic , Thomas Lucas , Fabien Baradel , Philippe Weinzaepfel , Gregory Rogez , Francesc Moreno-Noguer

Text-to-video models have demonstrated impressive capabilities in producing diverse and captivating video content, showcasing a notable advancement in generative AI. However, these models generally lack fine-grained control over motion…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Tuna Han Salih Meral , Hidir Yesiltepe , Connor Dunlop , Pinar Yanardag