中文
相关论文

相关论文: GeoViSTA: Geospatial Vision-Tabular Transformer fo…

200 篇论文

To mimic human vision with the way of recognizing the diverse and open world, foundation vision models are much critical. While recent techniques of self-supervised learning show the promising potentiality of this mission, we argue that…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Zhiming Qian

Multivariate geostatistics is based on modelling all covariances between all possible combinations of two or more variables at any sets of locations in a continuously indexed domain. Multivariate spatial covariance models need to be built…

统计方法学 · 统计学 2016-10-10 Noel Cressie , Andrew Zammit-Mangion

Vision transformers have recently emerged as an effective alternative to convolutional networks for action recognition. However, vision transformers still struggle with geometric variations prevalent in video data. This paper proposes a…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Jinhui Ye , Jiaming Zhou , Hui Xiong , Junwei Liang

Autonomous robot operation in unstructured environments is often underpinned by spatial understanding through vision. Systems composed of multiple concurrently operating robots additionally require access to frequent, accurate and reliable…

机器人学 · 计算机科学 2024-10-17 Jan Blumenkamp , Steven Morad , Jennifer Gielis , Amanda Prorok

How should we integrate representations from complementary sensors for autonomous driving? Geometry-based fusion has shown promise for perception (e.g. object detection, motion forecasting). However, in the context of end-to-end driving, we…

计算机视觉与模式识别 · 计算机科学 2022-06-01 Kashyap Chitta , Aditya Prakash , Bernhard Jaeger , Zehao Yu , Katrin Renz , Andreas Geiger

Data representation in non-Euclidean spaces has proven effective for capturing hierarchical and complex relationships in real-world datasets. Hyperbolic spaces, in particular, provide efficient embeddings for hierarchical structures. This…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Jacob Fein-Ashley , Ethan Feng , Minh Pham

Cross-view geo-localization is the problem of estimating the position and orientation (latitude, longitude and azimuth angle) of a camera at ground level given a large-scale database of geo-tagged aerial (e.g., satellite) images. Existing…

计算机视觉与模式识别 · 计算机科学 2020-05-11 Yujiao Shi , Xin Yu , Dylan Campbell , Hongdong Li

Unsupervised multimodal change detection is a practical and challenging topic that can play an important role in time-sensitive emergency applications. To address the challenge that multimodal remote sensing images cannot be directly…

计算机视觉与模式识别 · 计算机科学 2023-02-08 Hongruixuan Chen , Naoto Yokoya , Chen Wu , Bo Du

Capturing human mobility is essential for modeling how people interact with and move through physical spaces, reflecting social behavior, access to resources, and dynamic spatial patterns. To support scalable and transferable analysis…

人工智能 · 计算机科学 2025-06-18 Mohammad Hashemi , Andreas Zufle

Extracting compact, physically interpretable representations from high-dimensional scientific data is a persistent challenge due to the complex, nonlinear structures inherent in physical systems. We propose a Gaussian Mixture Variational…

机器学习 · 计算机科学 2025-12-01 Tiffany Fan , Murray Cutforth , Marta D'Elia , Alexandre Cortiella , Alireza Doostan , Eric Darve

Multi-modality spatio-temporal (MoST) data extends spatio-temporal (ST) data by incorporating multiple modalities, which is prevalent in monitoring systems, encompassing diverse traffic demands and air quality assessments. Despite…

机器学习 · 计算机科学 2024-05-07 Jiewen Deng , Renhe Jiang , Jiaqi Zhang , Xuan Song

Learning medical visual representations directly from paired radiology reports has become an emerging topic in representation learning. However, existing medical image-text joint learning methods are limited by instance or local supervision…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Fuying Wang , Yuyin Zhou , Shujun Wang , Varut Vardhanabhuti , Lequan Yu

Vision-centric Bird's Eye View (BEV) perception holds considerable promise for autonomous driving. Recent studies have prioritized efficiency or accuracy enhancements, yet the issue of domain shift has been overlooked, leading to…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Rongyu Zhang , Jiaming Liu , Xiaoqi Li , Xiaowei Chi , Dan Wang , Li Du , Yuan Du , Shanghang Zhang

We propose a task-agnostic framework for multimodal fusion of time series and single timestamp images, enabling cross-modal generation and robust downstream performance. Our approach explores deterministic and learned strategies for time…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Gianfranco Basile , Johannes Jakubik , Benedikt Blumenstiel , Thomas Brunschwiler , Juan Bernabe Moreno

3D-aware visual pretraining has proven effective in improving the performance of downstream robotic manipulation tasks. However, existing methods are constrained to Euclidean embedding spaces, whose flat geometry limits their ability to…

机器人学 · 计算机科学 2026-03-13 Jin Yang , Ping Wei , Yixin Chen , Nanning Zheng

Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAPE}, a closed-loop framework that combines generative visual…

机器人学 · 计算机科学 2025-09-03 Chuye Zhang , Xiaoxiong Zhang , Wei Pan , Linfang Zheng , Wei Zhang

Recent advancements in graph neural networks (GNNs) have significantly enhanced the prediction of material properties by modeling crystal structures as graphs. However, GNNs often struggle to capture global structural characteristics, such…

机器学习 · 计算机科学 2025-08-11 Jaewan Lee , Changyoung Park , Hongjun Yang , Sungbin Lim , Woohyung Lim , Sehui Han

For robots to robustly understand and interact with the physical world, it is highly beneficial to have a comprehensive representation - modelling geometry, physics, and visual observations - that informs perception, planning, and control…

机器人学 · 计算机科学 2024-06-18 Jad Abou-Chakra , Krishan Rana , Feras Dayoub , Niko Sünderhauf

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete…

机器学习 · 计算机科学 2026-05-13 Camillo Maria Caruso , Valerio Guarrasi , Paolo Soda
‹ 上一页 1 8 9 10 下一页 ›