English
Related papers

Related papers: Lang2Motion: Bridging Language and Motion through …

200 papers

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts describing the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Yuming Jiang , Shuai Yang , Tong Liang Koh , Wayne Wu , Chen Change Loy , Ziwei Liu

We introduce language-driven image generation, the task of generating an image visualizing the semantic contents of a word embedding, e.g., given the word embedding of grasshopper, we generate a natural image of a grasshopper. We implement…

Computer Vision and Pattern Recognition · Computer Science 2015-11-24 Angeliki Lazaridou , Dat Tien Nguyen , Raffaella Bernardi , Marco Baroni

Natural language is one of the most intuitive ways to express human intent. However, translating instructions and commands towards robotic motion generation and deployment in the real world is far from being an easy task. The challenge of…

Robotics · Computer Science 2022-09-20 Arthur Bucker , Luis Figueredo , Sami Haddadin , Ashish Kapoor , Shuang Ma , Sai Vemprala , Rogerio Bonatti

Foundation models have made significant strides in various applications, including text-to-image generation, panoptic segmentation, and natural language processing. This paper presents Instruct2Act, a framework that utilizes Large Language…

Robotics · Computer Science 2023-05-25 Siyuan Huang , Zhengkai Jiang , Hao Dong , Yu Qiao , Peng Gao , Hongsheng Li

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

We propose an attention-based networks for transferring motions between arbitrary objects. Given a source image(s) and a driving video, our networks animate the subject in the source images according to the motion in the driving video. In…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Subin Jeon , Seonghyeon Nam , Seoung Wug Oh , Seon Joo Kim

Automating pallet handling in outdoor logistics and construction environments remains challenging due to unstructured scenes, variable pallet configurations, and changing environmental conditions. In this paper, we present Lang2Lift, an…

Robotics · Computer Science 2026-02-26 Huy Hoang Nguyen , Johannes Huemer , Markus Murschitz , Tobias Glueck , Minh Nhat Vu , Andreas Kugi

Language models have demonstrated impressive ability in context understanding and generative performance. Inspired by the recent success of language foundation models, in this paper, we propose LMTraj (Language-based Multimodal Trajectory…

Computation and Language · Computer Science 2024-03-28 Inhwan Bae , Junoh Lee , Hae-Gon Jeon

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

We propose Point2Act, which directly retrieves the 3D action point relevant to a contextually described task, leveraging Multimodal Large Language Models (MLLMs). Foundation models opened the possibility for generalist robots that can…

Robotics · Computer Science 2026-03-05 Sang Min Kim , Hyeongjun Heo , Junho Kim , Yonghyeon Lee , Young Min Kim

Recently, human motion analysis has experienced great improvement due to inspiring generative models such as the denoising diffusion model and large language model. While the existing approaches mainly focus on generating motions with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Yiming Wu , Wei Ji , Kecheng Zheng , Zicheng Wang , Dong Xu

Text-to-video retrieval essentially aims to train models to align visual content with textual descriptions accurately. Due to the impressive general multimodal knowledge demonstrated by image-text pretrained models such as CLIP, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Yili Li , Gang Xiong , Gaopeng Gou , Xiangyan Qu , Jiamin Zhuang , Zhen Li , Junzheng Shi

Despite an exciting new wave of multimodal machine learning models, current approaches still struggle to interpret the complex contextual relationships between the different modalities present in videos. Going beyond existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Laura Hanu , Anita L. Verő , James Thewlis

The connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Zoltán Á. Milacski , Koichiro Niinuma , Ryosuke Kawamura , Fernando de la Torre , László A. Jeni

We address language-conditioned robotic manipulation using flow-based trajectory generation, which enables training on human and web videos of object manipulation and requires only minimal embodiment-specific data. This task is challenging,…

Robotics · Computer Science 2026-03-27 Motonari Kambara , Koki Seno , Tomoya Kaichi , Yanan Wang , Komei Sugiura

To represent motions from a mechanical point of view, this paper explores motion embedding using the motion taxonomy. With this taxonomy, manipulations can be described and represented as binary strings called motion codes. Motion codes…

Robotics · Computer Science 2020-07-15 David Paulius , Nicholas Eales , Yu Sun

Giving machines the ability to imagine possible new objects or scenes from linguistic descriptions and produce their realistic renderings is arguably one of the most challenging problems in computer vision. Recent advances in deep…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Levent Karacan , Tolga Kerimoğlu , İsmail İnan , Tolga Birdal , Erkut Erdem , Aykut Erdem

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Thomas Ressler-Antal , Frank Fundel , Malek Ben Alaya , Stefan Andreas Baumann , Felix Krause , Ming Gui , Björn Ommer

Videos are more informative than images because they capture the dynamics of the scene. By representing motion in videos, we can capture dynamic activities. In this work, we introduce GPT-4 generated motion descriptions that capture…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Chinmaya Devaraj , Cornelia Fermuller , Yiannis Aloimonos
‹ Prev 1 8 9 10 Next ›