English
Related papers

Related papers: DecoDINO: 3D Human-Scene Contact Prediction with S…

200 papers

Discriminative deep learning approaches have shown impressive results for problems where human-labeled ground truth is plentiful, but what about tasks where labels are difficult or impossible to obtain? This paper tackles one such problem:…

Computer Vision and Pattern Recognition · Computer Science 2016-04-20 Tinghui Zhou , Philipp Krähenbühl , Mathieu Aubry , Qixing Huang , Alexei A. Efros

Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guidance, yet these objectives are often treated as separate, disjoint goals. In this paper, we…

Robotics · Computer Science 2026-05-14 Han Yi Shin , Heeju Ko , Jaewon Mun , Qixing Huang , Jaehyeok Lee , Sung June Kim , Honglak Lee , Sujin Jang , Sangpil Kim

Recently, vision-language pre-training shows great potential in open-vocabulary object detection, where detectors trained on base classes are devised for detecting new classes. The class text embedding is firstly generated by feeding…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Yu Du , Fangyun Wei , Zihe Zhang , Miaojing Shi , Yue Gao , Guoqi Li

Recently, the DETR framework has emerged as the dominant approach for human--object interaction (HOI) research. In particular, two-stage transformer-based HOI detectors are amongst the most performant and training-efficient approaches.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Frederic Z. Zhang , Yuhui Yuan , Dylan Campbell , Zhuoyao Zhong , Stephen Gould

Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Fanjie Kong , Yanbei Chen , Jiarui Cai , Davide Modolo

We introduce a method for decentralized person re-identification in robot swarms that leverages natural language as the primary representational modality. Unlike traditional approaches that rely on opaque visual embeddings --…

Robotics · Computer Science 2026-01-21 Miquel Kegeleirs , Lorenzo Garattoni , Gianpiero Francesca , Mauro Birattari

Wearable robotics for lower-limb assistance have become a pivotal area of research, aiming to enhance mobility for individuals with physical impairments or augment the performance of able-bodied users. Accurate and adaptive control systems…

Collocated tactile sensing is a fundamental enabling technology for dexterous manipulation. However, deformable sensors introduce complex dynamics between the robot, grasped object, and environment that must be considered for fine…

Robotics · Computer Science 2022-09-28 Miquel Oller , Mireia Planas , Dmitry Berenson , Nima Fazeli

Comprehensive visual understanding requires detection frameworks that can effectively learn and utilize object interactions while analyzing objects individually. This is the main objective in Human-Object Interaction (HOI) detection task.…

Computer Vision and Pattern Recognition · Computer Science 2020-03-13 Oytun Ulutan , A S M Iftekhar , B. S. Manjunath

Learning 3D human-object interaction relation is pivotal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements, e.g., human contact, object affordance, and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yuhang Yang , Wei Zhai , Hongchen Luo , Yang Cao , Zheng-Jun Zha

Human-Object Interaction (HOI) detection aims to simultaneously localize human-object pairs and recognize their interactions. While recent two-stage approaches have made significant progress, they still face challenges due to incomplete…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Zhehao Li , Yucheng Qian , Chong Wang , Yinghao Lu , Zhihao Yang , Jiafei Wu

Long-horizon collaborative vision-language navigation (VLN) is critical for multi-robot systems to accomplish complex tasks beyond the capability of a single agent. CoNavBench takes a first step by introducing the first collaborative…

Robotics · Computer Science 2026-04-15 Sunyao Zhou , Yunzi Wu , Tianhang Wang , Xinhai Li , Guang Chen , Lizheng Liu , Chenjia Bai , Xuelong Li

Monocular vertex-level human-scene contact prediction is a fundamental capability for interactive systems such as assistive monitoring, embodied AI, and rehabilitation analysis. In this work, we study this task jointly with single-image 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xiaojian Lin , Yaomin Shen , Junyuan Ma , Yujie Sun , Chengqing Bu , Wenxin Zhang , Zongzheng Zhang , Hao Fei , Lei Jin , Hao Zhao

Predicting where people can walk in a scene is important for many tasks, including autonomous driving systems and human behavior analysis. Yet learning a computational model for this purpose is challenging due to semantic ambiguity and a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Jin Sun , Hadar Averbuch-Elor , Qianqian Wang , Noah Snavely

To understand the visual world, a machine must not only recognize individual object instances but also how they interact. Humans are often at the center of such interactions and detecting human-object interactions is an important practical…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Georgia Gkioxari , Ross Girshick , Piotr Dollár , Kaiming He

Equitable urban transportation applications require high-fidelity digital representations of the built environment: not just streets and sidewalks, but bike lanes, marked and unmarked crossings, curb ramps and cuts, obstructions, traffic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Bin Han , Yiwei Yang , Anat Caspi , Bill Howe

The development of algorithms to accurately decode neural information has long been a research focus in the field of neuroscience. Brain decoding typically involves training machine learning models to map neural data onto a preestablished…

Language-conditioned local navigation requires a robot to infer a nearby traversable target location from its current observation and an open-vocabulary, relational instruction. Existing vision-language spatial grounding methods usually…

Robotics · Computer Science 2026-03-11 Xinyu Gao , Gang Chen , Javier Alonso-Mora

One-shot transfer of dexterous grasps to novel scenes with object and context variations has been a challenging problem. While distilled feature fields from large vision models have enabled semantic correspondences across 3D scenes, their…

This paper explores a better prediction target for BERT pre-training of vision transformers. We observe that current prediction targets disagree with human perception judgment.This contradiction motivates us to learn a perceptual prediction…

Computer Vision and Pattern Recognition · Computer Science 2022-12-19 Xiaoyi Dong , Jianmin Bao , Ting Zhang , Dongdong Chen , Weiming Zhang , Lu Yuan , Dong Chen , Fang Wen , Nenghai Yu , Baining Guo
‹ Prev 1 4 5 6 7 8 10 Next ›