English
Related papers

Related papers: The ObjectFolder Benchmark: Multisensory Learning …

200 papers

The ability to understand the ways to interact with objects from visual cues, a.k.a. visual affordance, is essential to vision-guided robotic research. This involves categorizing, segmenting and reasoning of visual affordance. Relevant…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Shengheng Deng , Xun Xu , Chaozheng Wu , Ke Chen , Kui Jia

Standardized benchmarks are crucial for the majority of computer vision applications. Although leaderboards and ranking tables should not be over-claimed, benchmarks often provide the most objective measure of performance and are therefore…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Laura Leal-Taixé , Anton Milan , Konrad Schindler , Daniel Cremers , Ian Reid , Stefan Roth

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

The ability to associate touch with other modalities has huge implications for humans and computational systems. However, multimodal learning with touch remains challenging due to the expensive data collection process and non-standardized…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Fengyu Yang , Chao Feng , Ziyang Chen , Hyoungseob Park , Daniel Wang , Yiming Dou , Ziyao Zeng , Xien Chen , Rit Gangopadhyay , Andrew Owens , Alex Wong

Combining multiple sensors enables a robot to maximize its perceptual awareness of environments and enhance its robustness to external disturbance, crucial to robotic navigation. This paper proposes the FusionPortable benchmark, a complete…

The connection between visual input and tactile sensing is critical for object manipulation tasks such as grasping and pushing. In this work, we introduce the challenging task of estimating a set of tactile physical properties from visual…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Matthew Purri , Kristin Dana

The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and…

Computer Vision and Pattern Recognition · Computer Science 2022-07-08 Chuang Gan , Yi Gu , Siyuan Zhou , Jeremy Schwartz , Seth Alter , James Traer , Dan Gutfreund , Joshua B. Tenenbaum , Josh McDermott , Antonio Torralba

We propose a novel method for instance label segmentation of dense 3D voxel grids. We target volumetric scene representations, which have been acquired with depth sensors or multi-view stereo methods and which have been processed with…

Computer Vision and Pattern Recognition · Computer Science 2019-11-04 Jean Lahoud , Bernard Ghanem , Marc Pollefeys , Martin R. Oswald

Developing embodied agents in simulation has been a key research topic in recent years. Exciting new tasks, algorithms, and benchmarks have been developed in various simulators. However, most of them assume deaf agents in silent…

Robotics · Computer Science 2023-09-19 Ruohan Gao , Hao Li , Gokul Dharan , Zhuzhu Wang , Chengshu Li , Fei Xia , Silvio Savarese , Li Fei-Fei , Jiajun Wu

Object handover is a common human collaboration behavior that attracts attention from researchers in Robotics and Cognitive Science. Though visual perception plays an important role in the object handover task, the whole handover process…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Ruolin Ye , Wenqiang Xu , Zhendong Xue , Tutian Tang , Yanfeng Wang , Cewu Lu

Object detection is a fundamental visual recognition problem in computer vision and has been widely studied in the past decades. Visual object detection aims to find objects of certain target classes with precise localization in a given…

Computer Vision and Pattern Recognition · Computer Science 2019-08-13 Xiongwei Wu , Doyen Sahoo , Steven C. H. Hoi

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Weikun Peng , Denys Iliash , Manolis Savva

Deep learning techniques for point cloud data have demonstrated great potentials in solving classical problems in 3D computer vision such as 3D object classification and segmentation. Several recent 3D object classification methods have…

Computer Vision and Pattern Recognition · Computer Science 2019-08-20 Mikaela Angelina Uy , Quang-Hieu Pham , Binh-Son Hua , Duc Thanh Nguyen , Sai-Kit Yeung

Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. While recent multimodal embedding models excel at global image-text alignment, they often…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Shenghao Fu , Yukun Su , Fengyun Rao , Jing Lyu , Xiaohua Xie , Wei-Shi Zheng

The research presents an overhead view of 10 important objects and follows the general formatting requirements of the most popular machine learning task: digit recognition with MNIST. This dataset offers a public benchmark extracted from…

Computer Vision and Pattern Recognition · Computer Science 2021-02-09 David Noever , Samantha E. Miller Noever

Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. However, despite advancements, there remains a lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Shezheng Song , Chengxiang He , Shan Zhao , Chengyu Wang , Qian Wan , Tianwei Yan , Meng Wang

Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has focused primarily on recognition tasks such as meal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Sabab Ishraq , Aarushi Aarushi , Juncai Jiang , Chen Chen

Determining material properties from camera images can expand the ability to identify complex objects in indoor environments, which is valuable for consumer robotics applications. To support this, we introduce MatPredict, a dataset that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuzhen Chen , Hojun Son , Arpan Kusari

Despite increasing efforts on universal representations for visual recognition, few have addressed object detection. In this paper, we develop an effective and efficient universal object detection system that is capable of working on…

Computer Vision and Pattern Recognition · Computer Science 2019-07-09 Xudong Wang , Zhaowei Cai , Dashan Gao , Nuno Vasconcelos

Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and generalization…

‹ Prev 1 3 4 5 6 7 10 Next ›