English
Related papers

Related papers: Lang2Motion: Bridging Language and Motion through …

200 papers

This paper introduces a novel dataset construction pipeline that samples pairs of frames from videos and uses multimodal large language models (MLLMs) to generate editing instructions for training instruction-based image manipulation…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Mingdeng Cao , Xuaner Zhang , Yinqiang Zheng , Zhihao Xia

Segmenting objects with complex shapes, such as wires, bicycles, or structural grids, remains a significant challenge for current segmentation models, including the Segment Anything Model (SAM) and its high-quality variant SAM-HQ. These…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Luka Vetoshkin , Dmitry Yudin

Recent advances in deep learning have enabled the generation of videos from textual descriptions as well as the prediction of future sequences from input videos. Similarly, in human motion modeling, motions can be generated from text or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Masato Soga , Ryuki Takebayashi

Capturing and preserving motion semantics is essential to motion retargeting between animation characters. However, most of the previous works neglect the semantic information or rely on human-designed joint-level representations. Here, we…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Haodong Zhang , ZhiKe Chen , Haocheng Xu , Lei Hao , Xiaofei Wu , Songcen Xu , Zhensong Zhang , Yue Wang , Rong Xiong

Text-to-motion generation requires not only grounding local actions in language but also seamlessly blending these individual actions to synthesize diverse and realistic global motions. However, existing motion generation methods primarily…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Runyi Yu , Chang Liu , Xiangyang Ji , Li Yuan , Jie Chen

CLIP is a seminal multimodal model that maps images and text into a shared representation space through contrastive learning on billions of image-caption pairs. Inspired by the rapid progress of large language models (LLMs), we investigate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Weiquan Huang , Aoqi Wu , Yifan Yang , Xufang Luo , Yuqing Yang , Usman Naseem , Chunyu Wang , Chunyu Wang , Qi Dai , Xiyang Dai , Dongdong Chen , Chong Luo , Lili Qiu , Liang Hu

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zeyu Zhang , Yiran Wang , Wei Mao , Danning Li , Rui Zhao , Biao Wu , Zirui Song , Bohan Zhuang , Ian Reid , Richard Hartley

Multimodal Language Analysis is a demanding area of research, since it is associated with two requirements: combining different modalities and capturing temporal information. During the last years, several works have been proposed in the…

Computation and Language · Computer Science 2022-01-10 Panagiotis Koromilas , Theodoros Giannakopoulos

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, which have been shown to be effective at modeling language, each…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Xudong Lin , Gedas Bertasius , Jue Wang , Shih-Fu Chang , Devi Parikh , Lorenzo Torresani

The recent success of the CLIP model has shown its potential to be applied to a wide range of vision and language tasks. However this only establishes embedding space relationship of language to images, not to the video domain. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-04-11 Phani Krishna Uppala , Abhishek Bamotra , Shriti Priya , Vaidehi Joshi

We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using an image segmentation model and a shape completion model.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Zhao Mandi , Yijia Weng , Dominik Bauer , Shuran Song

We aim to control a robot to physically behave in the real world following any high-level language command like "cartwheel" or "kick". Although human motion datasets exist, this task remains particularly challenging since generative models…

Robotics · Computer Science 2024-05-21 Shusheng Xu , Huaijie Wang , Jiaxuan Gao , Yutao Ouyang , Chao Yu , Yi Wu

We consider the problem of segmenting objects in videos based on their motion and no other forms of supervision. Prior work has often approached this problem by using the principle of common fate, namely the fact that the motion of points…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Laurynas Karazija , Iro Laina , Christian Rupprecht , Andrea Vedaldi

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

In this work, we are dedicated to text-guided image generation and propose a novel framework, i.e., CLIP2GAN, by leveraging CLIP model and StyleGAN. The key idea of our CLIP2GAN is to bridge the output feature embedding space of CLIP and…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Yixuan Wang , Wengang Zhou , Jianmin Bao , Weilun Wang , Li Li , Houqiang Li

In this work, we focus on the problem of grounding language by training an agent to follow a set of natural language instructions and navigate to a target object in an environment. The agent receives visual information through raw pixels…

Computation and Language · Computer Science 2018-12-27 Akilesh B , Abhishek Sinha , Mausoom Sarkar , Balaji Krishnamurthy

Video motion transfer aims to generate a target video that inherits motion patterns from a source video while rendering new scenes. Existing training-free approaches focus on constructing motion guidance based on the intermediate outputs of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Zhen Wang , Youcan Xu , Jun Xiao , Long Chen

The Dense Trajectories concept is one of the most successful approaches in action recognition, suitable for scenarios involving a significant amount of motion. However, due to noise and background motion, many generated trajectories are…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Konstantinos Papadopoulos , Girum Demisse , Enjie Ghorbel , Michel Antunes , Djamila Aouada , Björn Ottersten

Deep-learning and large scale language-image training have produced image object detectors that generalise well to diverse environments and semantic classes. However, single-image object detectors trained on internet data are not optimally…

Robotics · Computer Science 2024-02-07 Nicolas Harvey Chapman , Feras Dayoub , Will Browne , Chris Lehnert