English
Related papers

Related papers: MUGL: Large Scale Multi Person Conditional Action …

200 papers

Decision-making in long-tail scenarios is pivotal to autonomous-driving development, and realistic and challenging simulations play a crucial role in testing safety-critical situations. However, existing open-source datasets lack systematic…

Robotics · Computer Science 2025-09-09 Chuancheng Zhang , Zhenhao Wang , Jiangcheng Wang , Kun Su , Qiang Lv , Bin Jiang , Kunkun Hao , Wenyu Wang

Developing the next generation of household robot helpers requires combining locomotion and interaction capabilities, which is generally referred to as mobile manipulation (MoMa). MoMa tasks are difficult due to the large action space of…

Robotics · Computer Science 2023-09-29 Jiaheng Hu , Peter Stone , Roberto Martín-Martín

We present an approach for 3D global human mesh recovery from monocular videos recorded with dynamic cameras. Our approach is robust to severe and long-term occlusions and tracks human bodies even when they go outside the camera's field of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Ye Yuan , Umar Iqbal , Pavlo Molchanov , Kris Kitani , Jan Kautz

3D Human motion generation is pivotal across film, animation, gaming, and embodied intelligence. Traditional 3D motion synthesis relies on costly motion capture, while recent work shows that 2D videos provide rich, temporally coherent…

Graphics · Computer Science 2026-05-20 Yi-Yang Zhang , Tengjiao Sun , Pengcheng Fang , Deng-Bao Wang , Xiaohao Cai , Min-Ling Zhang , Hansung Kim

For autonomous agents to successfully operate in real world, the ability to anticipate future motions of surrounding entities in the scene can greatly enhance their safety levels since potentially dangerous situations could be avoided in…

Machine Learning · Computer Science 2019-06-04 Yeping Hu , Wei Zhan , Liting Sun , Masayoshi Tomizuka

Large Language Models(LLMs) have revolutionized text generation and multimodal perception,but their capabilities in 3D content generation remain underexplored. Existing methods compromise by producing either low-resolution meshes or coarse…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Junming Huang , Chi Wang , Letian Li , Guangkai Xu , Donglin Huang , Hao Chen , Qiang Dai , Weiwei Xu

Current approaches to pose generation rely heavily on intermediate representations, either through two-stage pipelines with quantization or autoregressive models that accumulate errors during inference. This fundamental limitation leads to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yayuan Li , Filippos Bellos , Jason Corso

Micro-gesture recognition and behavior-based emotion prediction are both highly challenging tasks that require modeling subtle, fine-grained human behaviors, primarily leveraging video and skeletal pose data. In this work, we present two…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Arman Martirosyan , Shahane Tigranyan , Maria Razzhivina , Artak Aslanyan , Nazgul Salikhova , Ilya Makarov , Andrey Savchenko , Aram Avetisyan

Objective: Multimodal hand gesture recognition (HGR) systems can achieve higher recognition accuracy compared to unimodal HGR systems. However, acquiring multimodal gesture recognition data typically requires users to wear additional…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Wentao Wei , Linyan Ren

Understanding and generating multi-person interactions is a fundamental challenge with broad implications for robotics and social computing. While humans naturally coordinate in groups, modeling such interactions remains difficult due to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Vongani H. Maluleke , Kie Horiuchi , Lea Wilken , Evonne Ng , Jitendra Malik , Angjoo Kanazawa

In this work, we introduce an unconditional video generative model, InMoDeGAN, targeted to (a) generate high quality videos, as well as to (b) allow for interpretation of the latent space. For the latter, we place emphasis on interpreting…

Computer Vision and Pattern Recognition · Computer Science 2021-01-11 Yaohui Wang , Francois Bremond , Antitza Dantcheva

Recent advances in Large Language Models (LLMs) have shown impressive capabilities in various applications, yet LLMs face challenges such as limited context windows and difficulties in generalization. In this paper, we introduce a…

Neurons and Cognition · Quantitative Biology 2024-03-04 Jason Toy , Josh MacAdam , Phil Tabor

Lifelong sequence generation (LSG), a problem in continual learning, aims to continually train a model on a sequence of generation tasks to learn constantly emerging new generation patterns while avoiding the forgetting of previous…

Computation and Language · Computer Science 2023-11-23 Chengwei Qin , Chen Chen , Shafiq Joty

Large Language Models (LLMs) often suffer from hallucinations, which Retrieval-Augmented Generation (RAG) and GraphRAG mitigate by incorporating external knowledge and knowledge graphs (KGs). However, GraphRAG remains text-centric due to…

Artificial Intelligence · Computer Science 2026-03-11 Xueyao Wan , Hang Yu

Recently, unified multimodal models (UMMs) have made remarkable progress in integrating visual understanding and generation, demonstrating strong potential for complex text-to-image (T2I) tasks. Despite their theoretical promise, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jiadong Pan , Liang Li , Yuxin Peng , Yu-Ming Tang , Shuohuan Wang , Yu Sun , Hua Wu , Qingming Huang , Haifeng Wang

Human motion stylization aims to revise the style of an input motion while keeping its content unaltered. Unlike existing works that operate directly in pose space, we leverage the latent space of pretrained autoencoders as a more…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Chuan Guo , Yuxuan Mu , Xinxin Zuo , Peng Dai , Youliang Yan , Juwei Lu , Li Cheng

Motion synthesis in a dynamic environment has been a long-standing problem for character animation. Methods using motion capture data tend to scale poorly in complex environments because of their larger capturing and labeling requirement.…

Machine Learning · Computer Science 2021-01-06 Ying-Sheng Luo , Jonathan Hans Soeseno , Trista Pei-Chun Chen , Wei-Chao Chen

Multimodal large language models (MLLMs) have emerged as pivotal tools in enhancing human-computer interaction. In this paper we focus on the application of MLLMs in the field of graphical user interface (GUI) elements structuring, where…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yi Xu , Yesheng Zhang , Jiajia Liu , Jingdong Chen

As a unique biometric that can be perceived at a distance, gait has broad applications in person authentication, social security, and so on. Existing gait recognition methods suffer from changes in viewpoint and clothing and barely consider…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Jingqi Li , Jiaqi Gao , Yuzhen Zhang , Hongming Shan , Junping Zhang

Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tasks like visual question answering (VQA), image captioning,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Junxiao Xue , Quan Deng , Fei Yu , Yanhao Wang , Jun Wang , Yuehua Li