English
Related papers

Related papers: MonSTeR: a Unified Model for Motion, Scene, Text R…

200 papers

Localizing text instances in natural scenes is regarded as a fundamental challenge in computer vision. Nevertheless, owing to the extremely varied aspect ratios and scales of text instances in real scenes, most conventional text detectors…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Jingyang Lin , Yingwei Pan , Rongfeng Lai , Xuehang Yang , Hongyang Chao , Ting Yao

Scene text recognition (STR) from high-resolution (HR) images has been significantly successful, however text reading on low-resolution (LR) images is still challenging due to insufficient visual information. Therefore, recently many scene…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Minyi Zhao , Yang Wang , Jihong Guan , Shuigeng Zhou

We introduce HuMoR: a 3D Human Motion Model for Robust Estimation of temporal pose and shape. Though substantial progress has been made in estimating 3D human motion and shape from dynamic observations, recovering plausible pose sequences…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Davis Rempe , Tolga Birdal , Aaron Hertzmann , Jimei Yang , Srinath Sridhar , Leonidas J. Guibas

Walking as a form of active travel is essential in promoting sustainable transport. It is thus crucial to accurately predict pedestrian crossing intention and avoid collisions, especially with the advent of autonomous and advanced…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Chia-Yen Chiang , Yasmin Fathy , Gregory Slabaugh , Mona Jaber

This paper reports on a data-driven, interaction-aware motion prediction approach for pedestrians in environments cluttered with static obstacles. When navigating in such workspaces shared with humans, robots need accurate motion…

Robotics · Computer Science 2018-02-27 Mark Pfeiffer , Giuseppe Paolo , Hannes Sommer , Juan Nieto , Roland Siegwart , Cesar Cadena

Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applications such as immersive interaction and content creation. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xiaoxuan Ma , Jiashun Wang , Nicolas Ugrinovic , Yehonathan Litman , Kris Kitani

We present MST-MIXER - a novel video dialog model operating over a generic multi-modal state tracking scheme. Current models that claim to perform multi-modal state tracking fall short of two major aspects: (1) They either track only one…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Adnen Abdessaied , Lei Shi , Andreas Bulling

Object rearrangement is the problem of enabling a robot to identify the correct object placement in a complex environment. Prior work on object rearrangement has explored a diverse set of techniques for following user instructions to…

Robotics · Computer Science 2023-10-03 Kartik Ramachandruni , Max Zuo , Sonia Chernova

We introduce MonSter++, a geometric foundation model for multi-view depth estimation, unifying rectified stereo matching and unrectified multi-view stereo. Both tasks fundamentally recover metric depth from correspondence search and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Junda Cheng , Wenjing Liao , Zhipeng Cai , Longliang Liu , Gangwei Xu , Xianqi Wang , Yuzhou Wang , Zikang Yuan , Yong Deng , Jinliang Zang , Yangyang Shi , Jinhui Tang , Xin Yang

The ability to decompose complex multi-object scenes into meaningful abstractions like objects is fundamental to achieve higher-level cognition. Previous approaches for unsupervised object-oriented scene representation learning are either…

Machine Learning · Computer Science 2020-03-17 Zhixuan Lin , Yi-Fu Wu , Skand Vishwanath Peri , Weihao Sun , Gautam Singh , Fei Deng , Jindong Jiang , Sungjin Ahn

Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zesong Yang , Bangbang Yang , Wenqi Dong , Chenxuan Cao , Liyuan Cui , Yuewen Ma , Zhaopeng Cui , Hujun Bao

Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions commonly occur within diverse 3D scenes, which has prompted…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Ziyan Guo , Haoxuan Qu , Hossein Rahmani , Dewen Soh , Ping Hu , Qiuhong Ke , Jun Liu

High level understanding of sequential visual input is important for safe and stable autonomy, especially in localization and object detection. While traditional object classification and tracking approaches are specifically designed to…

Computer Vision and Pattern Recognition · Computer Science 2017-07-25 Mo Shan , Nikolay Atanasov

While most prior research has focused on improving the precision of multimodal trajectory predictions, the explicit modeling of multimodal behavioral intentions (e.g., yielding, overtaking) remains relatively underexplored. This paper…

Robotics · Computer Science 2026-02-19 Jiawei Sun , Xibin Yue , Jiahui Li , Tianle Shen , Chengran Yuan , Shuo Sun , Sheng Guo , Quanyun Zhou , Marcelo H Ang

Scene Text Recognition (STR), the task of recognizing text against complex image backgrounds, is an active area of research. Current state-of-the-art (SOTA) methods still struggle to recognize text written in arbitrary shapes. In this…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Ron Litman , Oron Anschel , Shahar Tsiper , Roee Litman , Shai Mazor , R. Manmatha

We study the problem of human action recognition using motion capture (MoCap) sequences. Unlike existing techniques that take multiple manual steps to derive standardized skeleton representations as model input, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Xiaoyu Zhu , Po-Yao Huang , Junwei Liang , Celso M. de Melo , Alexander Hauptmann

To achieve high coverage of target boxes, a normal strategy of conventional one-stage anchor-based detectors is to utilize multiple priors at each spatial position, especially in scene text detection tasks. In this work, we present a simple…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Linjie Deng , Yanxiang Gong , Xinchen Lu , Yi Lin , Zheng Ma , Mei Xie

Human motion and behaviour in crowded spaces is influenced by several factors, such as the dynamics of other moving agents in the scene, as well as the static elements that might be perceived as points of attraction or obstacles. In this…

Computer Vision and Pattern Recognition · Computer Science 2017-05-09 Federico Bartoli , Giuseppe Lisanti , Lamberto Ballan , Alberto Del Bimbo

As wearable sensing becomes increasingly pervasive, a key challenge remains: how can we generate natural language summaries from raw physiological signals such as actigraphy - minute-level movement data collected via accelerometers? In this…

Machine Learning · Computer Science 2025-12-29 Aiwei Zhang , Arvind Pillai , Andrew Campbell , Nicholas C. Jacobson

Traditional control and planning for robotic manipulation heavily rely on precise physical models and predefined action sequences. While effective in structured environments, such approaches often fail in real-world scenarios due to…

Robotics · Computer Science 2025-08-08 Jin Wang , Weijie Wang , Boyuan Deng , Heng Zhang , Rui Dai , Nikos Tsagarakis