中文
相关论文

相关论文: 2.5D Visual Relationship Detection

200 篇论文

To autonomously navigate and plan interactions in real-world environments, robots require the ability to robustly perceive and map complex, unstructured surrounding scenes. Besides building an internal representation of the observed scene…

机器人学 · 计算机科学 2021-05-18 Margarita Grinvald , Fadri Furrer , Tonci Novkovic , Jen Jen Chung , Cesar Cadena , Roland Siegwart , Juan Nieto

Vision and language tasks such as Visual Relation Detection and Visual Question Answering benefit from semantic features that afford proper grounding of language. The 3D depth of objects depicted in 2D images is one such feature. However it…

计算机视觉与模式识别 · 计算机科学 2021-01-13 Stefan Cassar , Adrian Muscat , Dylan Seychell

Visual Attention Models (VAMs) predict the location of an image or video regions that are most likely to attract human attention. Although saliency detection is well explored for 2D image and video content, there are only few attempts made…

图像与视频处理 · 电气工程与系统科学 2018-03-14 Amin Banitalebi-Dehkordi , Eleni Nasiopoulos , Mahsa T. Pourazad , Panos Nasiopoulos

Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently ill-posed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abundant 3D labels,…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Zihua Liu , Hiroki Sakuma , Masatoshi Okutomi

We propose Visual Query Detection (VQD), a new visual grounding task. In VQD, a system is guided by natural language to localize a variable number of objects in an image. VQD is related to visual referring expression recognition, where the…

计算机视觉与模式识别 · 计算机科学 2019-04-15 Manoj Acharya , Karan Jariwala , Christopher Kanan

In autonomous driving community, numerous benchmarks have been established to assist the tasks of 3D/2D object detection, stereo vision, semantic/instance segmentation. However, the more meaningful dynamic evolution of the surrounding…

计算机视觉与模式识别 · 计算机科学 2019-03-18 Jianru Xue , Jianwu Fang , Tao Li , Bohua Zhang , Pu Zhang , Zhen Ye , Jian Dou

Semantic segmentation of drone images is critical for various aerial vision tasks as it provides essential semantic details to understand scenes on the ground. Ensuring high accuracy of semantic segmentation models for drones requires…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Wenxiao Cai , Ke Jin , Jinyan Hou , Cong Guo , Letian Wu , Wankou Yang

Learning-based visual data compression and analysis have attracted great interest from both academia and industry recently. More training as well as testing datasets, especially good quality video datasets are highly desirable for related…

图像与视频处理 · 电气工程与系统科学 2021-05-14 Xiaozhong Xu , Shan Liu , Zeqiang Li

Remote sensing object detection has made significant progress, but most studies still focus on closed-set detection, limiting generalization across diverse datasets. Open-vocabulary object detection (OVD) provides a solution by leveraging…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Ziyue Huang , Yongchao Feng , Shuai Yang , Ziqi Liu , Qingjie Liu , Yunhong Wang

Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges…

计算机视觉与模式识别 · 计算机科学 2018-09-21 Fabien Baradel , Natalia Neverova , Christian Wolf , Julien Mille , Greg Mori

Image-to-text tasks, such as open-ended image captioning and controllable image description, have received extensive attention for decades. Here, we further advance this line of work by presenting Visual Spatial Description (VSD), a new…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Yu Zhao , Jianguo Wei , Zhichao Lin , Yueheng Sun , Meishan Zhang , Min Zhang

Modern perception systems of autonomous vehicles are known to be sensitive to occlusions and lack the capability of long perceiving range. It has been one of the key bottlenecks that prevents Level 5 autonomy. Recent research has…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Runsheng Xu , Xin Xia , Jinlong Li , Hanzhao Li , Shuo Zhang , Zhengzhong Tu , Zonglin Meng , Hao Xiang , Xiaoyu Dong , Rui Song , Hongkai Yu , Bolei Zhou , Jiaqi Ma

3D visual grounding aims to localize the target object in a 3D point cloud by a free-form language description. Typically, the sentences describing the target object tend to provide information about its relative relation between other…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Zehan Wang , Haifeng Huang , Yang Zhao , Linjun Li , Xize Cheng , Yichen Zhu , Aoxiong Yin , Zhou Zhao

Human interaction recognition is a challenging problem in computer vision and has been researched over the years due to its important applications. With the development of deep models for the human pose estimation problem, this work aims to…

计算机视觉与模式识别 · 计算机科学 2016-12-14 Marcel Sheeny de Moraes , Sankha Mukherjee , Neil M Robertson

Multi-view 3D object detection is a fundamental task in autonomous driving perception, where achieving a balance between detection accuracy and computational efficiency remains crucial. Sparse query-based 3D detectors efficiently aggregate…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Di Wu , Feng Yang , Wenhui Zhao , Jinwen Yu , Pan Liao , Benlian Xu , Dingwen Zhang

Multi-view Detection (MVD) is highly effective for occlusion reasoning in a crowded environment. While recent works using deep learning have made significant advances in the field, they have overlooked the generalization aspect, which makes…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Jeet Vora , Swetanjal Dutta , Kanishk Jain , Shyamgopal Karthik , Vineet Gandhi

Intelligent robots require object-level scene understanding to reason about possible tasks and interactions with the environment. Moreover, many perception tasks such as scene reconstruction, image retrieval, or place recognition can…

计算机视觉与模式识别 · 计算机科学 2023-05-05 Cathrin Elich , Iro Armeni , Martin R. Oswald , Marc Pollefeys , Joerg Stueckler

Visual saliency detection model simulates the human visual system to perceive the scene, and has been widely used in many vision tasks. With the acquisition technology development, more comprehensive information, such as depth cue,…

计算机视觉与模式识别 · 计算机科学 2019-09-04 Runmin Cong , Jianjun Lei , Huazhu Fu , Ming-Ming Cheng , Weisi Lin , Qingming Huang

Humans can easily deduce the relative pose of a previously unseen object, without labeling or training, given only a single query-reference image pair. This is arguably achieved by incorporating i) 3D/2.5D shape perception from a single…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Yuan Gao , Yajing Luo , Junhong Wang , Kui Jia , Gui-Song Xia

Fine-grained visual recognition is to classify objects with visually similar appearances into subcategories, which has made great progress with the development of deep CNNs. However, handling subtle differences between different…

计算机视觉与模式识别 · 计算机科学 2022-12-29 Yifan Zhao , Jia Li , Xiaowu Chen , Yonghong Tian