中文
相关论文

相关论文: KeypointNet: A Large-scale 3D Keypoint Dataset Agg…

200 篇论文

City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings, that can be…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Taiki Miyanishi , Fumiya Kitamori , Shuhei Kurita , Jungdae Lee , Motoaki Kawanabe , Nakamasa Inoue

MVImgNet is a large-scale dataset that contains multi-view images of ~220k real-world objects in 238 classes. As a counterpart of ImageNet, it introduces 3D visual signals via multi-view shooting, making a soft bridge between 2D and 3D…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Xiaoguang Han , Yushuang Wu , Luyue Shi , Haolin Liu , Hongjie Liao , Lingteng Qiu , Weihao Yuan , Xiaodong Gu , Zilong Dong , Shuguang Cui

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Yunnan Wang , Kecheng Zheng , Jianyuan Wang , Minghao Chen , David Novotny , Christian Rupprecht , Yinghao Xu , Xing Zhu , Wenjun Zeng , Xin Jin , Yujun Shen

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings,…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Patrick Rim , Kevin Harris , Braden Copple , Shangchen Han , Xu Xie , Ivan Shugurov , Sizhe An , He Wen , Alex Wong , Tomas Hodan , Kun He

In deep learning area, large-scale image datasets bring a breakthrough in the success of object recognition and retrieval. Nowadays, as the embodiment of innovation, the diversity of the industrial goods is significantly larger, in which…

计算机视觉与模式识别 · 计算机科学 2021-06-24 Fangyuan Lei , Da Huang , Jianjian Jiang , Ruijun Ma , Senhong Wang , Jiangzhong Cao , Yusen Lin , Qingyun Dai

In real-life scenarios, humans seek out objects in the 3D world to fulfill their daily needs or intentions. This inspires us to introduce 3D intention grounding, a new task in 3D object detection employing RGB-D, based on human intention,…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Weitai Kang , Mengxue Qu , Jyoti Kini , Yunchao Wei , Mubarak Shah , Yan Yan

We present a new public dataset with a focus on simulating robotic vision tasks in everyday indoor environments using real imagery. The dataset includes 20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely captured…

计算机视觉与模式识别 · 计算机科学 2017-03-07 Phil Ammirato , Patrick Poirson , Eunbyung Park , Jana Kosecka , Alexander C. Berg

We propose a method for annotating videos of complex multi-object scenes with a globally-consistent 3D representation of the objects. We annotate each object with a CAD model from a database, and place it in the 3D coordinate frame of the…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Kevis-Kokitsi Maninis , Stefan Popov , Matthias Nießner , Vittorio Ferrari

Query-based 3D object detection methods using multi-view images often struggle to efficiently leverage dynamic multi-scale information, e.g., the relationship between the object features and the geometric of the queries are not sufficiently…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Mingxi Pang , Dingheng Wang , Zekun Li , Zhenping Sun , Bo Wang , Zhihang Wang , Zhao-Xu Yang

Understanding charts requires models to jointly reason over geometric visual patterns, structured numerical data, and natural language -- a capability where current vision-language models (VLMs) remain limited. We introduce ChartNet, a…

With the increasing global popularity of self-driving cars, there is an immediate need for challenging real-world datasets for benchmarking and training various computer vision tasks such as 3D object detection. Existing datasets either…

计算机视觉与模式识别 · 计算机科学 2019-09-18 Quang-Hieu Pham , Pierre Sevestre , Ramanpreet Singh Pahwa , Huijing Zhan , Chun Ho Pang , Yuda Chen , Armin Mustafa , Vijay Chandrasekhar , Jie Lin

Estimating the 3D pose of desktop objects is crucial for applications such as robotic manipulation. Many existing approaches to this problem require a depth map of the object for both training and prediction, which restricts them to opaque,…

计算机视觉与模式识别 · 计算机科学 2020-05-20 Xingyu Liu , Rico Jonschkowski , Anelia Angelova , Kurt Konolige

Recently, there has been growing interest in developing learning-based methods to detect and utilize salient semi-global or global structures, such as junctions, lines, planes, cuboids, smooth surfaces, and all types of symmetries, for 3D…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Jia Zheng , Junfei Zhang , Jing Li , Rui Tang , Shenghua Gao , Zihan Zhou

Availability of a few, large-size, annotated datasets, like ImageNet, Pascal VOC and COCO, has lead deep learning to revolutionize computer vision research by achieving astonishing results in several vision tasks.We argue that new tools to…

计算机视觉与模式识别 · 计算机科学 2020-10-27 Pierluigi Zama Ramirez , Claudio Paternesi , Luca De Luigi , Luigi Lella , Daniele De Gregorio , Luigi Di Stefano

For the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other. Recently, however, the community has realized that…

计算机视觉与模式识别 · 计算机科学 2022-06-06 Yiyi Liao , Jun Xie , Andreas Geiger

Millions of people around the world have low or no vision. Assistive software applications have been developed for a variety of day-to-day tasks, including optical character recognition, scene identification, person recognition, and…

计算机视觉与模式识别 · 计算机科学 2022-04-11 Felipe Oviedo , Srinivas Vinnakota , Eugene Seleznev , Hemant Malhotra , Saqib Shaikh , Juan Lavista Ferres

Autonomous bin picking poses significant challenges to vision-driven robotic systems given the complexity of the problem, ranging from various sensor modalities, to highly entangled object layouts, to diverse item properties and gripper…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Maximilian Gilles , Yuhao Chen , Tim Robin Winter , E. Zhixuan Zeng , Alexander Wong

We introduce a new dataset for graphical object detection in business documents, more specifically annual reports. This dataset, IIIT-AR-13k, is created by manually annotating the bounding boxes of graphical or page objects in publicly…

计算机视觉与模式识别 · 计算机科学 2020-08-07 Ajoy Mondal , Peter Lipps , C. V. Jawahar

Monocular 3D object detection is well-known to be a challenging vision task due to the loss of depth information; attempts to recover depth using separate image-only approaches lead to unstable and noisy depth estimates, harming 3D…

计算机视觉与模式识别 · 计算机科学 2019-05-15 Ivan Barabanau , Alexey Artemov , Evgeny Burnaev , Vyacheslav Murashkin

Humans possess the cognitive ability to comprehend scenes in a compositional manner. To empower AI systems with similar capabilities, object-centric learning aims to acquire representations of individual objects from visual scenes without…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Yinxuan Huang , Tonglin Chen , Zhimeng Shen , Jinghao Huang , Bin Li , Xiangyang Xue