中文
相关论文

相关论文: CrowdMoGen: Zero-Shot Text-Driven Collective Motio…

200 篇论文

Predicting group behavior, how individuals coordinate, communicate, and interact during collaborative tasks, is essential for designing systems that can support team performance through real-time prediction and realistic simulation of…

人机交互 · 计算机科学 2026-04-13 Diana Romero , Xin Gao , Daniel Khalkhali , Salma Elmalaki

Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizing compound actions from texts, which integrate multiple concurrent actions into…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Yue Jiang , Mingyu Yang , Liuyuxin Yang , Yang Xu , Bingxin Yun , Yuhe Zhang

Human motion synthesis is a fundamental task in computer animation. Despite recent progress in this field utilizing deep learning and motion capture data, existing methods are always limited to specific motion categories, environments, and…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Zhikai Zhang , Yitang Li , Haofeng Huang , Mingxian Lin , Li Yi

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Wei Chen , Lin Li , Yongqi Yang , Bin Wen , Fan Yang , Tingting Gao , Yu Wu , Long Chen

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, how to effectively…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Xiaoyan Cong , Zekun Li , Zhiyang Dou , Hongyu Li , Omid Taheri , Chuan Guo , Abhay Mittal , Sizhe An , Taku Komura , Wojciech Matusik , Michael J. Black , Srinath Sridhar

Co-manipulation requires multiple humans to synchronize their motions with a shared object while ensuring reasonable interactions, maintaining natural poses, and preserving stable states. However, most existing motion generation approaches…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Jiahao Xu , Xiaohan Yuan , Xingchen Wu , Chongyang Xu , Kun Li , Buzhen Huang

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation by combining LLM and diffusion models, the state-of-the-art in each task, respectively. Existing approaches rely on spatial visual…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Kaihang Pan , Wang Lin , Zhongqi Yue , Tenglong Ao , Liyu Jia , Wei Zhao , Juncheng Li , Siliang Tang , Hanwang Zhang

Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck for large-scale deployment. To address this challenge, some works use LLMs and VLMs for…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Hao Shao , Letian Wang , Yang Zhou , Yuxuan Hu , Zhuofan Zong , Steven L. Waslander , Wei Zhan , Hongsheng Li

We have recently seen tremendous progress in diffusion advances for generating realistic human motions. Yet, they largely disregard the multi-human interactions. In this paper, we present InterGen, an effective diffusion-based approach that…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Han Liang , Wenqian Zhang , Wenxuan Li , Jingyi Yu , Lan Xu

Motion simulation, prediction and planning are foundational tasks in autonomous driving, each essential for modeling and reasoning about dynamic traffic scenarios. While often addressed in isolation due to their differing objectives, such…

机器人学 · 计算机科学 2026-02-03 Nan Song , Junzhe Jiang , Jingyu Li , Xiatian Zhu , Li Zhang

Achieving high-fidelity and temporally smooth 3D human motion generation remains a challenge, particularly within resource-constrained environments. We introduce FlowMotion, a novel method leveraging Conditional Flow Matching (CFM).…

图形学 · 计算机科学 2025-04-28 Manolo Canales Cuba , Vinícius do Carmo Melício , João Paulo Gois

With the growing demand for short videos and personalized content, automated Video Log (Vlog) generation has become a key direction in multimodal content creation. Existing methods mostly rely on predefined scripts, lacking dynamism and…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xiaolu Hou , Bing Ma , Jiaxiang Cheng , Xuhua Ren , Kai Yu , Wenyue Li , Tianxiang Zheng , Qinglin Lu

Generating 3D human motions from textual descriptions is an important research problem with broad applications in video games, virtual reality, and augmented reality. Recent methods align the textual description with human motion at the…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Bowen Dang , Lin Wu , Xiaohang Yang , Zheng Yuan , Zhixiang Chen

Text-to-motion (T2M) generation aims to control the behavior of a target character via textual descriptions. Leveraging text-motion paired datasets, existing T2M models have achieved impressive performance in generating high-quality motions…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Jiakun Zheng , Ting Xiao , Shiqin Cao , Xinran Li , Zhe Wang , Chenjia Bai

Generating realistic and controllable traffic scenes from natural language can greatly enhance the development and evaluation of autonomous driving systems. However, this task poses unique challenges: (1) grounding free-form text into…

机器人学 · 计算机科学 2026-03-27 Bo-Kai Ruan , Hao-Tang Tsui , Yung-Hui Li , Hong-Han Shuai

Text-to-video (T2V) synthesis has gained increasing attention in the community, in which the recently emerged diffusion models (DMs) have promisingly shown stronger performance than the past approaches. While existing state-of-the-art DMs…

人工智能 · 计算机科学 2024-03-20 Hao Fei , Shengqiong Wu , Wei Ji , Hanwang Zhang , Tat-Seng Chua

This paper addresses the problem of detecting coherent motions in crowd scenes and presents its two applications in crowd scene understanding: semantic region detection and recurrent activity mining. It processes input motion fields (e.g.,…

计算机视觉与模式识别 · 计算机科学 2016-04-20 Weiyao Lin , Yang Mi , Weiyue Wang , Jianxin Wu , Jingdong Wang , Tao Mei

We present CoMet, a novel approach for computing a group's cohesion and using that to improve a robot's navigation in crowded scenes. Our approach uses a novel cohesion-metric that builds on prior work in social psychology. We compute this…

Recent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in videos generated by any…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Changgu Chen , Junwei Shu , Gaoqi He , Changbo Wang , Yang Li

This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Muhammad Islam , Tao Huang , Euijoon Ahn , Usman Naseem