English
Related papers

Related papers: Encoder-Free Human Motion Understanding via Struct…

200 papers

We develop a technique for generating smooth and accurate 3D human pose and motion estimates from RGB video sequences. Our method, which we call Motion Estimation via Variational Autoencoder (MEVA), decomposes a temporal sequence of human…

Computer Vision and Pattern Recognition · Computer Science 2020-10-07 Zhengyi Luo , S. Alireza Golestaneh , Kris M. Kitani

This paper advances motion agents empowered by large language models (LLMs) toward autonomous navigation in dynamic and cluttered environments, significantly surpassing first and recent seminal but limited studies on LLM's spatial…

Artificial Intelligence · Computer Science 2025-06-06 Yubo Zhao , Qi Wu , Yifan Wang , Yu-Wing Tai , Chi-Keung Tang

Estimating 3D human pose and shape from a single image is highly under-constrained. To address this ambiguity, we propose a novel prior, namely kinematic dictionary, which explicitly regularizes the solution space of relative 3D rotations…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Ze Ma , Yifan Yao , Pan Ji , Chao Ma

It is well-established that more data generally improves AI model performance. However, data collection can be challenging for certain tasks due to the rarity of occurrences or high costs. These challenges are evident in our use case, where…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Martin Thißen , Thi Ngoc Diep Tran , Barbara Esteve Ratsch , Ben Joel Schönbein , Ute Trapp , Beate Egner , Romana Piat , Elke Hergenröther

One key challenge in multi-document summarization is to capture the relations among input documents that distinguish between single document summarization (SDS) and multi-document summarization (MDS). Few existing MDS works address this…

Computation and Language · Computer Science 2022-09-14 Congbo Ma , Wei Emma Zhang , Pitawelayalage Dasun Dileepa Pitawela , Yutong Qu , Haojie Zhuang , Hu Wang

The growing exploration of Large Language Models (LLM) and Vision-Language Models (VLM) has opened avenues for enhancing the effectiveness of reinforcement learning (RL). However, existing LLM-based RL methods often focus on the guidance of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Wentao Wang , Chunyang Liu , Kehua Sheng , Bo Zhang , Yan Wang

Predicting future human motion plays a significant role in human-machine interactions for various real-life applications. A unified formulation and multi-order modeling are two critical perspectives for analyzing and representing human…

Computer Vision and Pattern Recognition · Computer Science 2021-12-30 Xiaoli Liu , Jianqin Yin , Huaping Liu , Jun Liu

Gait silhouettes, which can be encoded into binary gait codes, are widely adopted to representing motion patterns of pedestrian. Recent approaches commonly leverage visual backbones to encode gait silhouettes, achieving successful…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Ruiyi Zhan , Guozhen Peng , Canyu Chen , Jian Lei , Annan Li

Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis framework that take…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Gong Jingyu , Tong Kunkun , Chen Zhuoran , Yuan Chuanhan , Chen Mingang , Zhang Zhizhong , Tan Xin , Xie Yuan

Personal robots assisting humans must perform complex manipulation tasks that are typically difficult to specify in traditional motion planning pipelines, where multiple objectives must be met and the high-level context be taken into…

Robotics · Computer Science 2019-03-21 Hejia Zhang , Eric Heiden , Stefanos Nikolaidis , Joseph J. Lim , Gaurav S. Sukhatme

Large Language Models (LLMs) have demonstrated remarkable performance across various domains, including healthcare. However, their ability to effectively represent structured non-textual data, such as the alphanumeric medical codes used in…

Visual encoding constitutes the basis of large multimodal models (LMMs) in understanding the visual world. Conventional LMMs process images in fixed sizes and limited resolutions, while recent explorations in this direction are limited in…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Ruyi Xu , Yuan Yao , Zonghao Guo , Junbo Cui , Zanlin Ni , Chunjiang Ge , Tat-Seng Chua , Zhiyuan Liu , Maosong Sun , Gao Huang

Our aim is to develop a unified model for sign language understanding, that performs sign language translation (SLT) and sign-subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Youngjoon Jang , Liliane Momeni , Zifan Jiang , Joon Son Chung , Gül Varol , Andrew Zisserman

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Luchuan Song , Pinxin Liu , Haiyang Liu , Zhenchao Jin , Yolo Yunlong Tang , Zichong Xu , Susan Liang , Jing Bi , Jason J Corso , Chenliang Xu

Motion planning in complex scenarios is the core challenge in autonomous driving. Conventional methods apply predefined rules or learn from driving data to plan the future trajectory. Recent methods seek the knowledge preserved in large…

Robotics · Computer Science 2024-06-12 Ruijun Zhang , Xianda Guo , Wenzhao Zheng , Chenming Zhang , Kurt Keutzer , Long Chen

Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior work has primarily focused on textual or visual-textual…

Computation and Language · Computer Science 2025-09-22 Zhu Li , Xiyuan Gao , Yuqing Zhang , Shekhar Nayak , Matt Coler

Skeleton-based human action recognition has achieved remarkable progress in recent years. However, most existing GCN-based methods rely on short-range motion topologies, which not only struggle to capture long-range joint dependencies and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Ruosi Wang , Fangwei Zuo , Lei Li , Zhaoqiang Xia

We propose a method to create document representations that reflect their internal structure. We modify Tree-LSTMs to hierarchically merge basic elements such as words and sentences into blocks of increasing complexity. Our Structure…

Computation and Language · Computer Science 2019-10-08 Khalil Mrini , Claudiu Musat , Michael Baeriswyl , Martin Jaggi

In clinical practice, segmenting specific lesions based on the needs of physicians can significantly enhance diagnostic accuracy and treatment efficiency. However, conventional lesion segmentation models lack the flexibility to distinguish…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Shuyi Ouyang , Jinyang Zhang , Xiangye Lin , Xilai Wang , Qingqing Chen , Yen-Wei Chen , Lanfen Lin

Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, notably in connecting ideas and adhering to logical rules to solve problems. These models have evolved to accommodate various data modalities, including sound…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Enis Berk Çoban , Michael I. Mandel , Johanna Devaney
‹ Prev 1 8 9 10 Next ›