English
Related papers

Related papers: TextIM: Part-aware Interactive Motion Synthesis fr…

200 papers

Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typically follow either a two-stage paradigm that combines kinematic motion generation with…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Bin Li , Ruichi Zhang , Han Liang , Jingyan Zhang , Juze Zhang , Xin Chen , Jingya Wang

Tactility provides crucial support and enhancement for the perception and interaction capabilities of both humans and robots. Nevertheless, the multimodal research related to touch primarily focuses on visual and tactile modalities, with…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Ning Cheng , You Li , Jing Gao , Bin Fang , Jinan Xu , Wenjuan Han

Generating immersive 3D scenes from texts is a core task in computer vision, crucial for applications in virtual reality and game development. Despite the promise of leveraging 2D diffusion priors, existing methods suffer from spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Jisheng Chu , Wenrui Li , Rui Zhao , Wangmeng Zuo , Shifeng Chen , Xiaopeng Fan

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Yuming Jiang , Shuai Yang , Tong Liang Koh , Wayne Wu , Chen Change Loy , Ziwei Liu

With diffusion transformer (DiT) excelling in video generation, its use in specific tasks has drawn increasing attention. However, adapting DiT for pose-guided human image animation faces two core challenges: (a) existing U-Net-based pose…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Haoyu Zhao , Zhongang Qi , Cong Wang , Qingping Zheng , Guansong Lu , Fei Chen , Hang Xu , Zuxuan Wu

We introduce InteractVLM, a novel method to estimate 3D contact points on human bodies and objects from single in-the-wild images, enabling accurate human-object joint reconstruction in 3D. This is challenging due to occlusions, depth…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Sai Kumar Dwivedi , Dimitrije Antić , Shashank Tripathi , Omid Taheri , Cordelia Schmid , Michael J. Black , Dimitrios Tzionas

A good co-speech motion generation cannot be achieved without a careful integration of common rhythmic motion and rare yet essential semantic motion. In this work, we propose SemTalk for holistic co-speech motion generation with frame-level…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Xiangyue Zhang , Jianfang Li , Jiaxu Zhang , Ziqiang Dang , Jianqiang Ren , Liefeng Bo , Zhigang Tu

Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textual semantics. However, most existing works study these two problems separately, without…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Jidong Kuang , Hongsong Wang , Jie Gui

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Despite impressive advances, existing approaches are often…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yuan Wang , Di Huang , Yaqi Zhang , Wanli Ouyang , Jile Jiao , Xuetao Feng , Yan Zhou , Pengfei Wan , Shixiang Tang , Dan Xu

Generating realistic human motions from textual descriptions has undergone significant advancements. However, existing methods often overlook specific body part movements and their timing. In this paper, we address this issue by enriching…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Bizhu Wu , Jinheng Xie , Meidan Ding , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Lightweight, controllable, and physically plausible human motion synthesis is crucial for animation, virtual reality, robotics, and human-computer interaction applications. Existing methods often compromise between computational efficiency,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Arvin Tashakori , Arash Tashakori , Gongbo Yang , Z. Jane Wang , Peyman Servati

In the realm of motion generation, the creation of long-duration, high-quality motion sequences remains a significant challenge. This paper presents our groundbreaking work on "Infinite Motion", a novel approach that leverages long text to…

Computer Vision and Pattern Recognition · Computer Science 2024-07-15 Mengtian Li , Chengshuo Zhai , Shengxiang Yao , Zhifeng Xie , Keyu Chen , Yu-Gang Jiang

We introduce the Multi-Motion Discrete Diffusion Models (M2D2M), a novel approach for human motion generation from textual descriptions of multiple actions, utilizing the strengths of discrete diffusion models. This approach adeptly…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Seunggeun Chi , Hyung-gun Chi , Hengbo Ma , Nakul Agarwal , Faizan Siddiqui , Karthik Ramani , Kwonjoon Lee

Effective human-robot interaction requires emotionally rich multimodal expressions, yet most humanoid robots lack coordinated speech, facial expressions, and gestures. Meanwhile, real-world deployment demands on-device solutions that can…

Robotics · Computer Science 2026-02-10 Songhua Yang , Xuetao Li , Xuanye Fei , Mengde Li , Miao Li

Novel-View Human Action Synthesis aims to synthesize the movement of a body from a virtual viewpoint, given a video from a real viewpoint. We present a novel 3D reasoning to synthesize the target viewpoint. We first estimate the 3D mesh of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-09 Mohamed Ilyes Lakhal , Davide Boscaini , Fabio Poiesi , Oswald Lanz , Andrea Cavallaro

This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores. Our work will impact various retail applications that need better customer…

Computation and Language · Computer Science 2024-09-20 Hikaru Asano , Ryo Yonetani , Taiki Sekii , Hiroki Ouchi

In this paper, we propose composable part-based manipulation (CPM), a novel approach that leverages object-part decomposition and part-part correspondences to improve learning and generalization of robotic manipulation skills. By…

Robotics · Computer Science 2024-05-10 Weiyu Liu , Jiayuan Mao , Joy Hsu , Tucker Hermans , Animesh Garg , Jiajun Wu

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Bizhu Wu , Jinheng Xie , Keming Shen , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence…

Computer Vision and Pattern Recognition · Computer Science 2019-11-01 Yitian Yuan , Lin Ma , Jingwen Wang , Wei Liu , Wenwu Zhu

Generating natural and expressive human motions from textual descriptions is challenging due to the complexity of coordinating full-body dynamics and capturing nuanced motion patterns over extended sequences that accurately reflect the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Seong-Eun Hong , Soobin Lim , Juyeong Hwang , Minwook Chang , Hyeongyeop Kang