中文
相关论文

相关论文: GeoVista: Visually Grounded Active Perception for …

200 篇论文

Recent advances in Vision-Language-Action (VLA) models have enabled robotic agents to integrate multimodal understanding with action execution. However, our empirical analysis reveals that current VLAs struggle to allocate visual attention…

Remote sensing imagery offers rich spectral data across extensive areas for Earth observation. Many attempts have been made to leverage these data with transfer learning to develop scalable alternatives for estimating socio-economic…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Fan Yang , Sahoko Ishida , Mengyan Zhang , Daniel Jenson , Swapnil Mishra , Jhonathan Navott , Seth Flaxman

Global geolocation, which seeks to predict the geographical location of images captured anywhere in the world, is one of the most challenging tasks in the field of computer vision. In this paper, we introduce an innovative interactive…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Zhiyang Dou , Zipeng Wang , Xumeng Han , Guorong Li , Zhipei Huang , Zhenjun Han

Considering the accelerated development of Unmanned Aerial Vehicles (UAVs) applications in both industrial and research scenarios, there is an increasing need for localizing these aerial systems in non-urban environments, using GNSS-Free,…

机器人学 · 计算机科学 2022-10-19 Marius-Mihail Gurgu , Jorge Peña Queralta , Tomi Westerlund

Efficient processing of high-resolution images is crucial for real-world vision-language applications. However, existing Large Vision-Language Models (LVLMs) incur substantial computational overhead due to the large number of vision tokens.…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Jewon Lee , Wooksu Shin , Seungmin Yang , Ki-Ung Song , DongUk Lim , Jaeyeon Kim , Tae-Ho Kim , Bo-Kyeong Kim

Visual Grounding (VG) aims to locate the most relevant region in an image, based on a flexible natural language query but not a pre-defined label, thus it can be a more useful technique than object detection in practice. Most…

计算机视觉与模式识别 · 计算机科学 2019-03-19 Chaorui Deng , Qi Wu , Guanghui Xu , Zhuliang Yu , Yanwu Xu , Kui Jia , Mingkui Tan

Deep learning models are essential for scene classification, change detection, land cover segmentation, and other remote sensing image understanding tasks. Most backbones of existing remote sensing deep learning models are typically…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Ziyue Huang , Mingming Zhang , Yuan Gong , Qingjie Liu , Yunhong Wang

Remote sensing question answering (RS-QA) often requires more than direct semantic prediction, especially in large-scale forest scenes where ecological analysis involves multi-step filtering, numerical aggregation, neighborhood reasoning,…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Zihang Cheng , Duanchu Wang , Cheng Li , Jing Huang , Huanzhao Fu , Di Wang

Tenant evictions threaten housing stability and are a major concern for many cities. An open question concerns whether data-driven methods enhance outreach programs that target at-risk tenants to mitigate their risk of eviction. We propose…

In this work we study the problem of exploring surfaces and building compact 3D representations of the environment surrounding a robot through active perception. We propose an online probabilistic framework that merges visual and tactile…

机器人学 · 计算机科学 2018-02-14 Sergio Caccamo , Yasemin Bekiroglu , Carl Henrik Ek , Danica Kragic

Cross-View Geo-Localization tackles the challenge of image geo-localization in GNSS-denied environments, including disaster response scenarios, urban canyons, and dense forests, by matching street-view query images with geo-tagged…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Panwang Xia , Lei Yu , Yi Wan , Qiong Wu , Peiqi Chen , Liheng Zhong , Yongxiang Yao , Dong Wei , Xinyi Liu , Lixiang Ru , Yingying Zhang , Jiangwei Lao , Jingdong Chen , Ming Yang , Yongjun Zhang

In recent years, Virtual Reality (VR) Head-Mounted Displays (HMD) have been used to provide an immersive, first-person view in real-time for the remote-control of Unmanned Ground Vehicles (UGV). One critical issue is that it is challenging…

人机交互 · 计算机科学 2022-01-11 Yiming Luo , Jialin Wang , Rongkai Shi , Hai-Ning Liang , Shan Luo

Multimodal large language models (MLLMs) have made rapid progress in recent years, yet continue to struggle with low-level visual perception (LLVP) -- particularly the ability to accurately describe the geometric details of an image. This…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Jiarui Zhang , Ollie Liu , Tianyu Yu , Jinyi Hu , Willie Neiswanger

Modern cameras are equipped with a wide array of sensors that enable recording the geospatial context of an image. Taking advantage of this, we explore depth estimation under the assumption that the camera is geocalibrated, a problem we…

计算机视觉与模式识别 · 计算机科学 2021-09-22 Scott Workman , Hunter Blanton

Universal Photometric Stereo is a promising approach for recovering surface normals without strict lighting assumptions. However, it struggles when multi-illumination cues are unreliable, such as under biased lighting or in shadows or…

计算机视觉与模式识别 · 计算机科学 2025-11-19 King-Man Tam , Satoshi Ikehata , Yuta Asano , Zhaoyi An , Rei Kawakami

We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable…

机器人学 · 计算机科学 2024-10-11 Naoki Wake , Atsushi Kanehira , Kazuhiro Sasabuchi , Jun Takamatsu , Katsushi Ikeuchi

Unsupervised pre-training strategies have proven to be highly effective in natural language processing and computer vision. Likewise, unsupervised reinforcement learning (RL) holds the promise of discovering a variety of potentially useful…

机器学习 · 计算机科学 2024-03-12 Seohong Park , Oleh Rybkin , Sergey Levine

In ground-view object change detection, the recently emerging mapless navigation has great potential to navigate a robot to objects distantly detected (e.g., books, cups, clothes) and acquire high-resolution object images, to identify their…

机器人学 · 计算机科学 2023-10-25 Kouki Terashima , Kanji Tanaka , Ryogo Yamamoto , Jonathan Tay Yu Liang

Active vision, also known as active perception, refers to the process of actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Muzhi Zhu , Hao Zhong , Canyu Zhao , Zongze Du , Zheng Huang , Mingyu Liu , Hao Chen , Cheng Zou , Jingdong Chen , Ming Yang , Chunhua Shen

Deep learning based object detection has achieved great success. However, these supervised learning methods are data-hungry and time-consuming. This restriction makes them unsuitable for limited data and urgent tasks, especially in the…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Tengfei Zhang , Yue Zhang , Xian Sun , Menglong Yan , Yaoling Wang , Kun Fu