English
Related papers

Related papers: RAMEN: Resolution-Adjustable Multimodal Encoder fo…

200 papers

The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on…

This paper presents RAVEN, a computationally efficient deep learning architecture for FMCW radar perception. The method processes raw ADC data in a chirp-wise streaming manner, preserves MIMO structure through independent receiver…

Signal Processing · Electrical Eng. & Systems 2026-04-07 Anuvab Sen , Mir Sayeed Mohammad , Saibal Mukhopadhyay

Utilizing the complementary strengths of wavelength-specific range or depth sensors is crucial for robust computer-assisted tasks such as autonomous driving. Despite this, there is still little research done at the intersection of optical…

Image and Video Processing · Electrical Eng. & Systems 2025-11-11 Vanessa Wirth , Johanna Bräunig , Nikolai Hofmann , Martin Vossiek , Tim Weyrich , Marc Stamminger

Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks. However, we observe that their feature representations are poorly aligned across different modalities. For instance, the feature embedding…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Rishabh Kabra , Maks Ovsjanikov , Drew A. Hudson , Ye Xia , Skanda Koppula , Andre Araujo , Joao Carreira , Niloy J. Mitra

Learning representations of nodes has been a crucial area of the graph machine learning research area. A well-defined node embedding model should reflect both node features and the graph structure in the final embedding. In the case of…

Machine Learning · Computer Science 2023-04-20 Kamil Tagowski , Piotr Bielak , Jakub Binkowski , Tomasz Kajdanowicz

The rapid evolution of Vision Language Models (VLMs) has catalyzed significant advancements in artificial intelligence, expanding research across various disciplines, including Earth Observation (EO). While VLMs have enhanced image…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Xizhe Xue , Guoting Wei , Hao Chen , Haokui Zhang , Feng Lin , Chunhua Shen , Xiao Xiang Zhu

Earth observation (EO) data features diverse sensing platforms with varying spectral bands, spatial resolutions, and sensing modalities. While most prior work has constrained inputs to fixed sensors, a new class of any-sensor foundation…

Machine Learning · Computer Science 2025-08-04 Leonard Waldmann , Ando Shah , Yi Wang , Nils Lehmann , Adam J. Stewart , Zhitong Xiong , Xiao Xiang Zhu , Stefan Bauer , John Chuang

Earth observation (EO) foundation models (FMs) are increasingly trained on multisensor data, spanning multispectral imagery (MSI), synthetic aperture radar (SAR), and derived geospatial layers, but hyperspectral imagery (HSI) remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Nassim Ait Ali Braham , Aaron Banze , Conrad M. Albrecht , Julien Mairal , Jocelyn Chanussot , Xiao Xiang Zhu

Earth Observation (EO) data are increasingly used in policy analysis by enabling granular estimation of conditional average treatment effects (CATE). However, a challenge in EO-based causal inference is determining the scale of the input…

Machine Learning · Statistics 2025-03-18 Fucheng Warren Zhu , Connor T. Jerzak , Adel Daoud

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack…

Machine Learning · Computer Science 2025-12-02 Jacob Thompson , Emiliano Garcia-Lopez , Yonatan Bisk

We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent…

Unsupervised multimodal change detection is pivotal for time-sensitive tasks and comprehensive multi-temporal Earth monitoring. In this study, we explore unsupervised multimodal change detection between two key remote sensing data sources:…

Image and Video Processing · Electrical Eng. & Systems 2024-01-18 Hongruixuan Chen , Jian Song , Naoto Yokoya

With the proliferation of mobile devices, the need for an efficient model to restore any degraded image has become increasingly significant and impactful. Traditional approaches typically involve training dedicated models for each specific…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Bin Ren , Eduard Zamfir , Zongwei Wu , Yawei Li , Yidi Li , Danda Pani Paudel , Radu Timofte , Ming-Hsuan Yang , Nicu Sebe

Existing multi-object image generation methods face difficulties in achieving precise alignment between localized image generation regions and their corresponding semantics based on language descriptions, frequently resulting in…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Yanfeng Li , Yue Sun , Keren Fu , Sio-Kei Im , Xiaoming Liu , Guangtao Zhai , Xiaohong Liu , Tao Tan

Weather prediction is a quintessential problem involving the forecasting of a complex, nonlinear, and chaotic high-dimensional dynamical system. This work introduces an efficient reduced-order modeling (ROM) framework for short-range…

Machine Learning · Computer Science 2025-11-18 Amirpasha Hedayat , Karthik Duraisamy

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geographic coverage, or…

Beyond the existing single-person and multiple-person human parsing tasks in static images, this paper makes the first attempt to investigate a more realistic video instance-level human parsing that simultaneously segments out each person…

Computer Vision and Pattern Recognition · Computer Science 2018-08-13 Qixian Zhou , Xiaodan Liang , Ke Gong , Liang Lin

A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar sensors, on the other hand, are becoming invaluable due to their robustness in variable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Kavin Chandrasekaran , Sorin Grigorescu , Gijs Dubbelman , Pavol Jancura

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by…

Machine Learning · Computer Science 2023-04-06 Alexandros Haliassos , Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Maja Pantic

Satellite-based remote sensing has revolutionised the way we address global challenges. Huge quantities of Earth Observation (EO) data are generated by satellite sensors daily, but processing these large datasets for use in ML pipelines is…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Matthew J Allen , Francisco Dorr , Joseph Alejandro Gallego Mejia , Laura Martínez-Ferrer , Anna Jungbluth , Freddie Kalaitzis , Raúl Ramos-Pollán