中文
相关论文

相关论文: ObjectFolder: A Dataset of Objects with Implicit V…

200 篇论文

This paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric representations of images. Compared to the most widely used…

计算机视觉与模式识别 · 计算机科学 2021-03-11 Pengchuan Zhang , Xiujun Li , Xiaowei Hu , Jianwei Yang , Lei Zhang , Lijuan Wang , Yejin Choi , Jianfeng Gao

Detecting oriented tiny objects, which are limited in appearance information yet prevalent in real-world applications, remains an intricate and under-explored problem. To address this, we systemically introduce a new dataset, benchmark, and…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Chang Xu , Ruixiang Zhang , Wen Yang , Haoran Zhu , Fang Xu , Jian Ding , Gui-Song Xia

Multi-view 3D object detection is becoming popular in autonomous driving due to its high effectiveness and low cost. Most of the current state-of-the-art detectors follow the query-based bird's-eye-view (BEV) paradigm, which benefits from…

计算机视觉与模式识别 · 计算机科学 2023-06-05 Zhangyang Qi , Jiaqi Wang , Xiaoyang Wu , Hengshuang Zhao

Augmented data storytelling enhances narrative delivery by integrating visualizations with physical environments and presenter actions. Existing systems predominantly rely on body gestures or speech to control visualizations, leaving…

人机交互 · 计算机科学 2025-08-08 Kentaro Takahira , Yue Yu , Takanori Fujiwara , Ryo Suzuki , Huamin Qu

For visually impaired people, it is highly difficult to make independent movement and safely move in both indoors and outdoors environment. Furthermore, these physically and visually challenges prevent them from in day-today live…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Heba Najm , Khirallah Elferjani , Alhaam Alariyibi

We tackle the task of scalable unsupervised object-centric representation learning on 3D scenes. Existing approaches to object-centric representation learning show limitations in generalizing to larger scenes as their learning processes…

计算机视觉与模式识别 · 计算机科学 2023-09-26 Tianyu Wang , Kee Siong Ng , Miaomiao Liu

In this paper, we address the problem of estimating the in-hand 6D pose of an object in contact with multiple vision-based tactile sensors. We reason on the possible spatial configurations of the sensors along the object surface.…

机器人学 · 计算机科学 2023-02-01 Gabriele M. Caddeo , Nicola A. Piga , Fabrizio Bottarel , Lorenzo Natale

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Zitian Tang , Shijie Wang , Junho Cho , Jaewook Yoo , Chen Sun

This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the ``Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D,…

Due to large variations in shape, appearance, and viewing conditions, object recognition is a key precursory challenge in the fields of object manipulation and robotic/AI visual reasoning in general. Recognizing object categories,…

计算机视觉与模式识别 · 计算机科学 2015-04-14 Haopeng Zhang , Tarek El-Gaaly , Ahmed Elgammal , Zhiguo Jiang

We present the DeepScores dataset with the goal of advancing the state-of-the-art in small objects recognition, and by placing the question of object recognition in the context of scene understanding. DeepScores contains high quality images…

计算机视觉与模式识别 · 计算机科学 2018-05-29 Lukas Tuggener , Ismail Elezi , Jürgen Schmidhuber , Marcello Pelillo , Thilo Stadelmann

We present ObjectBox, a novel single-stage anchor-free and highly generalizable object detection approach. As opposed to both existing anchor-based and anchor-free detectors, which are more biased toward specific object scales in their…

计算机视觉与模式识别 · 计算机科学 2022-07-15 Mohsen Zand , Ali Etemad , Michael Greenspan

People with visual impairments face numerous challenges when interacting with their environment. Our objective is to develop a device that facilitates communication between individuals with visual impairments and their surroundings. The…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Souayah Abdelkader , Mokretar Kraroubi Abderrahmene , Slimane Larabi

Recent advancements in deep learning, computer vision, and embodied AI have given rise to synthetic causal reasoning video datasets. These datasets facilitate the development of AI algorithms that can reason about physical interactions…

人工智能 · 计算机科学 2021-08-16 Jiafei Duan , Samson Yu Bai Jian , Cheston Tan

Object detection is a fundamental visual recognition problem in computer vision and has been widely studied in the past decades. Visual object detection aims to find objects of certain target classes with precise localization in a given…

计算机视觉与模式识别 · 计算机科学 2019-08-13 Xiongwei Wu , Doyen Sahoo , Steven C. H. Hoi

Coreset selection is a method for selecting a small, representative subset of an entire dataset. It has been primarily researched in image classification, assuming there is only one object per image. However, coreset selection for object…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Hojun Lee , Suyoung Kim , Junhoo Lee , Jaeyoung Yoo , Nojun Kwak

We present a new dataset with annotated eye movements. The dataset consists of over 800,000 gaze points recorded during a car ride in the real world and in the simulator. In total, the eye movements of 19 subjects were annotated. In this…

计算机视觉与模式识别 · 计算机科学 2021-01-13 Wolfgang Fuhl , Enkelejda Kasneci

Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Sreehari Rajan , Kunal Bhosikar , Charu Sharma

Visual control enables quadrotors to adaptively navigate using real-time sensory data, bridging perception with action. Yet, challenges persist, including generalization across scenarios, maintaining reliability, and ensuring real-time…

机器人学 · 计算机科学 2024-04-09 Alessandro Saviolo , Pratyaksh Rao , Vivek Radhakrishnan , Jiuhong Xiao , Giuseppe Loianno

Understanding and forecasting future scene states is critical for autonomous agents to plan and act effectively in complex environments. Object-centric models, with structured latent spaces, have shown promise in modeling object dynamics…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Angel Villar-Corrales , Gjergj Plepi , Sven Behnke