English
Related papers

Related papers: AnyMo: Scaling Any-Modality Conditional Motion Gen…

200 papers

We tackle the challenges of synthesizing versatile, physically simulated human motions for full-body object manipulation. Unlike prior methods that are focused on detailed motion tracking, trajectory following, or teleoperation, our…

Robotics · Computer Science 2025-12-12 Chen Tessler , Yifeng Jiang , Erwin Coumans , Zhengyi Luo , Gal Chechik , Xue Bin Peng

Fitting an underlying body model to 3D clothed human assets has been extensively studied, yet most approaches focus on either single-modal inputs such as point clouds or multi-view images alone, often requiring a known metric scale. This…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Zeyu Cai , Yuliang Xiu , Renke Wang , Zhijing Shao , Xiaoben Li , Siyuan Yu , Chao Xu , Yang Liu , Baigui Sun , Jian Yang , Zhenyu Zhang

Music-driven 3D dance generation offers significant creative potential, yet practical applications demand versatile and multimodal control. As the highly dynamic and complex human motion covering various styles and genres, dance generation…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Jinlu Zhang , Zixi Kang , Libin Liu , Jianlong Chang , Qi Tian , Feng Gao , Yizhou Wang

This paper proposes MotionVerse, a unified framework that harnesses the capabilities of Large Language Models (LLMs) to comprehend, generate, and edit human motion in both single-person and multi-person scenarios. To efficiently represent…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ruibing Hou , Mingshuang Luo , Hongyu Pan , Hong Chang , Shiguang Shan

The emergence of neural networks has revolutionized the field of motion synthesis. Yet, learning to unconditionally synthesize motions from a given distribution remains challenging, especially when the motions are highly diverse. In this…

Graphics · Computer Science 2022-12-20 Sigal Raab , Inbal Leibovitch , Peizhuo Li , Kfir Aberman , Olga Sorkine-Hornung , Daniel Cohen-Or

Inspired by cognitive theories, we introduce AnyHome, a framework that translates any text into well-structured and textured indoor scenes at a house-scale. By prompting Large Language Models (LLMs) with designed templates, our approach…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Rao Fu , Zehao Wen , Zichen Liu , Srinath Sridhar

Text-guided dynamic 3D character generation has advanced rapidly, yet producing high-quality motion that faithfully reflects rich textual descriptions remains challenging. Existing methods tend to generate limited sub-actions or incoherent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Miaowei Wang , Qingxuan Yan , Zhi Cao , Yayuan Li , Oisin Mac Aodha , Jason J. Corso , Amir Vaxman

In this report, we present HoloMotion-1, a humanoid motion foundation model for zero-shot whole-body motion tracking. A key innovation of HoloMotion-1 is to scale control-policy training with a large-scale hybrid motion corpus, where…

Robotics · Computer Science 2026-05-20 Maiyue Chen , Kaihui Wang , Bo Zhang , Xihan Ma , Zhiyuan Yang , Yi Ren , Qijun Huang , Zihao Zhu , Yucheng Wang , Zhizhong Su

With the rapid progress of large language models (LLMs), multimodal frameworks that unify understanding and generation have become promising, yet they face increasing complexity as the number of modalities and tasks grows. We observe that…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Bingfan Zhu , Biao Jiang , Sunyi Wang , Shixiang Tang , Tao Chen , Linjie Luo , Youyi Zheng , Xin Chen

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study…

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, how to effectively…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Xiaoyan Cong , Zekun Li , Zhiyang Dou , Hongyu Li , Omid Taheri , Chuan Guo , Abhay Mittal , Sizhe An , Taku Komura , Wojciech Matusik , Michael J. Black , Srinath Sridhar

Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Tao Tang , Enhui Ma , xia zhou , Letian Wang , Tianyi Yan , Xueyang Zhang , Kun Zhan , Peng Jia , XianPeng Lang , Jia-Wang Bian , Kaicheng Yu , Xiaodan Liang

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present SkyReels-Audio, a unified framework for synthesizing…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zhengcong Fei , Hao Jiang , Di Qiu , Baoxuan Gu , Youqiang Zhang , Jiahua Wang , Jialin Bai , Debang Li , Mingyuan Fan , Guibin Chen , Yahui Zhou

In the molecular domain, numerous studies have explored the use of multimodal large language models (LLMs) to construct a general-purpose, multi-task molecular model. However, these efforts are still far from achieving a truly universal…

Machine Learning · Computer Science 2025-10-31 Chengxin Hu , Hao Li , Yihe Yuan , Zezheng Song , Chenyang Zhao , Haixin Wang

This paper extends Cross Modal Generalization (CMG) to open-set environments by proposing the more challenging Open-set Cross Modal Generalization (OSCMG) task. This task evaluates multimodal unified representations in open-set conditions,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Hai Huang , Yan Xia , Shulei Wang , Hanting Wang , Minghui Fang , Shengpeng Ji , Sashuai Zhou , Tao Jin , Zhou Zhao

Hand-Object Interaction (HOI) generation has significant application potential. However, current 3D HOI motion generation approaches heavily rely on predefined 3D object models and lab-captured motion data, limiting generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Lingwei Dang , Ruizhi Shao , Hongwen Zhang , Wei Min , Yebin Liu , Qingyao Wu

Lightweight, controllable, and physically plausible human motion synthesis is crucial for animation, virtual reality, robotics, and human-computer interaction applications. Existing methods often compromise between computational efficiency,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Arvin Tashakori , Arash Tashakori , Gongbo Yang , Z. Jane Wang , Peyman Servati

Scaling up motion datasets is crucial to enhance motion generation capabilities. However, training on large-scale multi-source datasets introduces data heterogeneity challenges due to variations in motion content. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Junyu Shi , Lijiang Liu , Yong Sun , Zhiyuan Zhang , Jinni Zhou , Qiang Nie

We propose cross-modal attentive connections, a new dynamic and effective technique for multimodal representation learning from wearable data. Our solution can be integrated into any stage of the pipeline, i.e., after any convolutional…

Machine Learning · Computer Science 2022-06-10 Anubhav Bhatti , Behnam Behinaein , Paul Hungler , Ali Etemad