中文
相关论文

相关论文: GeoViT: A Versatile Vision Transformer Architectur…

200 篇论文

Hydrometric forecasting is crucial for managing water resources, flood prediction, and environmental protection. Water stations are interconnected, and this connectivity influences the measurements at other stations. However, the dynamic…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Naghmeh Shafiee Roudbari , Ursula Eicker , Charalambos Poullis , Zachary Patterson

This paper presents a vision transformer (ViT) based joint source and channel coding (JSCC) scheme for wireless image transmission over multiple-input multiple-output (MIMO) systems, called ViT-MIMO. The proposed ViT-MIMO architecture, in…

信息论 · 计算机科学 2022-10-28 Haotian Wu , Yulin Shao , Chenghong Bian , Krystian Mikolajczyk , Deniz Gündüz

Reasoning about visual relationships is central to how humans interpret the visual world. This task remains challenging for current deep learning algorithms since it requires addressing three key technical problems jointly: 1) identifying…

计算机视觉与模式识别 · 计算机科学 2022-06-14 Xiaojian Ma , Weili Nie , Zhiding Yu , Huaizu Jiang , Chaowei Xiao , Yuke Zhu , Song-Chun Zhu , Anima Anandkumar

Machine learning researchers strive to develop better and better algorithms to solve computer vision problems, such as image classification. In recent years, the classification of micro-Doppler spectrograms has also benefited from these…

信号处理 · 电气工程与系统科学 2025-12-02 Arkadiusz Czuba

Rapid environmental change and advances in data-driven analysis highlight the need not only to use computational tools, but also to foster understanding of the natural world and inspire creativity. Photosynthesis, the process that fuels…

In this work, we address the problem of cross-view geo-localization, which estimates the geospatial location of a street view image by matching it with a database of geo-tagged aerial images. The cross-view matching task is extremely…

计算机视觉与模式识别 · 计算机科学 2021-07-06 Hongji Yang , Xiufan Lu , Yingying Zhu

This paper introduces a novel visual analytics approach, DCPViz, to enable climate scientists to explore massive climate data interactively without requiring the upfront movement of massive data. Thus, climate scientists are afforded more…

Accurate global localization is critical for autonomous driving and robotics, but GNSS-based approaches often degrade due to occlusion and multipath effects. As an emerging alternative, cross-view pose estimation predicts the 3-DoF camera…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Juhye Park , Wooju Lee , Dasol Hong , Changki Sung , Youngwoo Seo , Dongwan Kang , Hyun Myung

Estimating the 6-degrees-of-freedom (6DoF) pose of a spacecraft from a single image is critical for autonomous operations like in-orbit servicing and space debris removal. Existing state-of-the-art methods often rely on iterative…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Pierre Ancey , Andrew Price , Saqib Javed , Mathieu Salzmann

The task of cross-view image geo-localization aims to determine the geo-location (GPS coordinates) of a query ground-view image by matching it with the GPS-tagged aerial (satellite) images in a reference dataset. Due to the dramatic changes…

计算机视觉与模式识别 · 计算机科学 2019-04-15 Bin Sun , Chen Chen , Yingying Zhu , Jianmin Jiang

Learning Bird's Eye View (BEV) representation from surrounding-view cameras is of great importance for autonomous driving. In this work, we propose a Geometry-guided Kernel Transformer (GKT), a novel 2D-to-BEV representation learning…

计算机视觉与模式识别 · 计算机科学 2022-06-10 Shaoyu Chen , Tianheng Cheng , Xinggang Wang , Wenming Meng , Qian Zhang , Wenyu Liu

Accurate Subseasonal-to-Seasonal (S2S) climate forecasting is pivotal for decision-making including agriculture planning and disaster preparedness but is known to be challenging due to its chaotic nature. Although recent data-driven models…

机器学习 · 计算机科学 2025-02-28 Yang Liu , Zinan Zheng , Jiashun Cheng , Fugee Tsung , Deli Zhao , Yu Rong , Jia Li

Cross-view geo-localization (CVGL), which matches an oblique drone view to a geo-referenced satellite tile, has emerged as a key alternative for autonomous drone navigation when GNSS signals are jammed, spoofed, or unavailable. Despite…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Chi-Nguyen Tran , Dao Sy Duy Minh , Huynh Trung Kiet , Nguyen Lam Phu Quy , Phu-Hoa Pham , Long Tran-Thanh

Vision Transformer (ViT) has emerged as a powerful architecture in the realm of modern computer vision. However, its application in certain imaging fields, such as microscopy and satellite imaging, presents unique challenges. In these…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Yujia Bao , Srinivasan Sivanandan , Theofanis Karaletsos

Recently, Transformers have emerged as the go-to architecture for both vision and language modeling tasks, but their computational efficiency is limited by the length of the input sequence. To address this, several efficient variants of…

计算机视觉与模式识别 · 计算机科学 2023-05-18 Hao Zheng , Jinbao Wang , Xiantong Zhen , Hong Chen , Jingkuan Song , Feng Zheng

Transformers are state-of-the-art deep learning models that are composed of stacked attention and point-wise, fully connected layers designed for handling sequential data. Transformers are not only ubiquitous throughout Natural Language…

计算机视觉与模式识别 · 计算机科学 2021-12-01 Onur Kara , Arijit Sehanobish , Hector H Corzo

Predicting typhoon intensity accurately across space and time is crucial for issuing timely disaster warnings and facilitating emergency response. This has vast potential for minimizing life losses and property damages as well as reducing…

计算机视觉与模式识别 · 计算机科学 2023-12-05 Huanxin Chen , Pengshuai Yin , Huichou Huang , Qingyao Wu , Ruirui Liu , Xiatian Zhu

Gravitational lensing offers a powerful probe into the properties of dark matter and is crucial to infer cosmological parameters. The Legacy Survey of Space and Time (LSST) is predicted to find O(10^5) gravitational lenses over the next…

计算机视觉与模式识别 · 计算机科学 2025-09-03 René Parlange , Juan C. Cuevas-Tello , Octavio Valenzuela , Omar de J. Cabrera-Rosas , Tomás Verdugo , Anupreeta More , Anton T. Jaelani

Reliable flood detection is critical for disaster management, yet classical deep learning models often struggle with the high-dimensional, nonlinear complexities inherent in remote sensing data. To mitigate these limitations, we introduced…

机器学习 · 计算机科学 2026-03-17 Soumyajit Maity , Behzad Ghanbarian

In this work we demonstrate the vulnerability of vision transformers (ViTs) to gradient-based inversion attacks. During this attack, the original data batch is reconstructed given model weights and the corresponding gradients. We introduce…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Ali Hatamizadeh , Hongxu Yin , Holger Roth , Wenqi Li , Jan Kautz , Daguang Xu , Pavlo Molchanov