中文
相关论文

相关论文: GAMa: Cross-view Video Geo-localization

200 篇论文

To perform outdoor visual navigation and search, a robot may leverage satellite imagery to generate visual priors. This can help inform high-level search strategies, even when such images lack sufficient resolution for target recognition.…

With the rapid growth of the low-altitude economy, UAVs have become crucial for measurement and tracking in patrol systems. However, in GNSS-denied areas, satellite-based localization methods are prone to failure. This paper presents a…

计算机视觉与模式识别 · 计算机科学 2025-11-05 Tao Liu , Kan Ren , Qian Chen

Unmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) presents significant challenges due to the view discrepancy between oblique UAV images and overhead satellite images. Existing methods heavily rely on the supervision of…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Haoyuan Li , Chang Xu , Wen Yang , Li Mi , Huai Yu , Haijian Zhang

In some scenarios, a single input image may not be enough to allow the object classification. In those cases, it is crucial to explore the complementary information extracted from images presenting the same object from multiple perspectives…

计算机视觉与模式识别 · 计算机科学 2022-05-24 Gabriel Machado , Keiller Nogueira , Matheus Barros Pereira , Jefersson Alex dos Santos

In this paper, we focus on video relocalization task, which uses a query video clip as input to retrieve a semantic relative video clip in another untrimmed long video. we find that in video relocalization datasets, there exists a…

计算机视觉与模式识别 · 计算机科学 2022-01-27 Yuan Zhou , Mingfei Wang , Ruolin Wang , Shuwei Huo

Feature matching is a crucial task in the field of computer vision, which involves finding correspondences between images. Previous studies achieve remarkable performance using learning-based feature comparison. However, the pervasive…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yesheng Zhang , Xu Zhao

Neural rendering methods can achieve near-photorealistic image synthesis of scenes from posed input images. However, when the images are imperfect, e.g., captured in very low-light conditions, state-of-the-art methods fail to reconstruct…

计算机视觉与模式识别 · 计算机科学 2024-07-12 Vinayak Gupta , Rongali Simhachala Venkata Girish , Mukund Varma T , Ayush Tewari , Kaushik Mitra

Existing large video-language models (LVLMs) struggle to comprehend long videos correctly due to limited context. To address this problem, fine-tuning long-context LVLMs and employing GPT-based agents have emerged as promising solutions.…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongdong Luo , Xiawu Zheng , Guilin Li , Shukang Yin , Haojia Lin , Chaoyou Fu , Jinfa Huang , Jiayi Ji , Fei Chao , Jiebo Luo , Rongrong Ji

Cross-modal retrieval aims to measure the content similarity between different types of data. The idea has been previously applied to visual, text, and speech data. In this paper, we present a novel cross-modal retrieval method specifically…

计算机视觉与模式识别 · 计算机科学 2020-05-05 Numan Khurshid , Talha Hanif , Mohbat Tharani , Murtaza Taj

Many problems can be viewed as forms of geospatial search aided by aerial imagery, with examples ranging from detecting poaching activity to human trafficking. We model this class of problems in a visual active search (VAS) framework, which…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Anindya Sarkar , Michael Lanier , Scott Alfeld , Jiarui Feng , Roman Garnett , Nathan Jacobs , Yevgeniy Vorobeychik

Existing video domain adaption (DA) methods need to store all temporal combinations of video frames or pair the source and target videos, which are memory cost expensive and can't scale up to long videos. To address these limitations, we…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Xinyue Hu , Lin Gu , Liangchen Liu , Ruijiang Li , Chang Su , Tatsuya Harada , Yingying Zhu

High Dynamic Range (HDR) content (i.e., images and videos) has a broad range of applications. However, capturing HDR content from real-world scenes is expensive and time-consuming. Therefore, the challenging task of reconstructing visually…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Hrishav Bakul Barua , Kalin Stefanov , KokSheik Wong , Abhinav Dhall , Ganesh Krishnasamy

Large Language Model (LLM)-based Automated Program Repair (APR) has shown strong potential on textual benchmarks, yet struggles in multimodal scenarios where bugs are reported with GUI screenshots. Existing methods typically convert images…

软件工程 · 计算机科学 2026-04-10 Zhuoyao Liu , Zhengran Zeng , Shu-Dong Huang , Yang Liu , Shikun Zhang , Wei Ye

Worldwide geolocalization aims to locate the precise location at the coordinate level of photos taken anywhere on the Earth. It is very challenging due to 1) the difficulty of capturing subtle location-aware visual semantics, and 2) the…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Pengyue Jia , Yiding Liu , Xiaopeng Li , Yuhao Wang , Yantong Du , Xiao Han , Xuetao Wei , Shuaiqiang Wang , Dawei Yin , Xiangyu Zhao

Cross-view geo-localization determines the location of a query image, captured by a drone or ground-based camera, by matching it to a geo-referenced satellite image. While traditional approaches focus on image-level localization, many…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Zheyang Huang , Jagannath Aryal , Saeid Nahavandi , Xuequan Lu , Chee Peng Lim , Lei Wei , Hailing Zhou

In this work, we tackle the challenging problem of unsupervised video domain adaptation (UVDA) for action recognition. We specifically focus on scenarios with a substantial domain gap, in contrast to existing works primarily deal with small…

计算机视觉与模式识别 · 计算机科学 2023-11-23 Hyogun Lee , Kyungho Bae , Seong Jong Ha , Yumin Ko , Gyeong-Moon Park , Jinwoo Choi

Visual re-localization means using a single image as input to estimate the camera's location and orientation relative to a pre-recorded environment. The highest-scoring methods are "structure based," and need the query camera's intrinsics…

计算机视觉与模式识别 · 计算机科学 2021-04-13 Mehmet Ozgur Turkoglu , Eric Brachmann , Konrad Schindler , Gabriel Brostow , Aron Monszpart

Compositing-aware object search aims to find the most compatible objects for compositing given a background image and a query bounding box. Previous works focus on learning compatibility between the foreground object and background, but…

计算机视觉与模式识别 · 计算机科学 2022-04-04 Sijie Zhu , Zhe Lin , Scott Cohen , Jason Kuen , Zhifei Zhang , Chen Chen

To track the target in a video, current visual trackers usually adopt greedy search for target object localization in each frame, that is, the candidate region with the maximum response score will be selected as the tracking result of each…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Xiao Wang , Zhe Chen , Bo Jiang , Jin Tang , Bin Luo , Dacheng Tao

Given a ground-level query image and a geo-referenced aerial image that covers the query's local surroundings, fine-grained cross-view localization aims to estimate the location of the ground camera inside the aerial image. Recent works…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Zimin Xia , Yujiao Shi , Hongdong Li , Julian F. P. Kooij