English
Related papers

Related papers: RAMEN: Resolution-Adjustable Multimodal Encoder fo…

200 papers

The increasing availability of Earth observation data offers unprecedented opportunities for large-scale environmental monitoring and analysis. However, these datasets are inherently heterogeneous, stemming from diverse sensors,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Georges Le Bellier , Nicolas Audebert

We address prevailing challenges of the brain-powered research, departing from the observation that the literature hardly recover accurate spatial information and require subject-specific models. To address these challenges, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Weihao Xia , Raoul de Charette , Cengiz Öztireli , Jing-Hao Xue

The advancement of remote sensing, including satellite systems, facilitates the continuous acquisition of remote sensing imagery globally, introducing novel challenges for achieving open-world tasks. Deployed models need to continuously…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Xiang Xiang , Zhuo Xu , Yao Deng , Qinhao Zhou , Yifan Liang , Ke Chen , Qingfang Zheng , Yaowei Wang , Xilin Chen , Wen Gao

Multimodal variational autoencoders have demonstrated their ability to learn the relationships between different modalities by mapping them into a latent representation. Their design and capacity to perform any-to-any conditional and…

Machine Learning · Computer Science 2025-02-04 Daniel Wesego , Pedram Rooshenas

Masked Autoencoders (MAEs) achieve impressive performance in image classification tasks, yet the internal representations they learn remain less understood. This work started as an attempt to understand the strong downstream classification…

Machine Learning · Computer Science 2026-02-04 Anika Shrivastava , Renu Rameshan , Samar Agnihotri

Computational surface modeling that underlies material recognition has transitioned from reflectance modeling using in-lab controlled radiometric measurements to image-based representations based on internet-mined single-view images…

Computer Vision and Pattern Recognition · Computer Science 2020-09-24 Jia Xue , Hang Zhang , Ko Nishino , Kristin J. Dana

Existing studies on bundle construction have relied merely on user feedback via bipartite graphs or enhanced item representations using semantic information. These approaches fail to capture elaborate relations hidden in real-world bundle…

Dark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET)…

Computer Vision and Pattern Recognition · Computer Science 2022-05-09 Ziteng Cui , Guo-Jun Qi , Lin Gu , Shaodi You , Zenghui Zhang , Tatsuya Harada

Recent work by Oberti et al, (Astron. Astrophys., 667, 48, 2022) argued and made a compelling case that classical astronomical adaptive optics (AO) tomography performance can be further enhanced by carefully designing and optically…

Instrumentation and Methods for Astrophysics · Physics 2025-12-22 Carlos M. Correia , Pierre Jouve , Jesse Cranney , Guido Agapito Cédric Taïssir Heritier

Multimodal learning has been a popular area of research, yet integrating electroencephalogram (EEG) data poses unique challenges due to its inherent variability and limited availability. In this paper, we introduce a novel multimodal…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Kang Yin , Hye-Bin Shin , Dan Li , Seong-Whan Lee

Synthetic Aperture Radar (SAR) and optical imagery provide complementary strengths that constitute the critical foundation for transcending single-modality constraints and facilitating cross-modal collaborative processing and intelligent…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Peihao Wu , Yongxiang Yao , Yi Wan , Wenfei Zhang , Ruipeng Zhao , Jiayuan Li , Yongjun Zhang

Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full modality integrity is often not applicable in reality and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-07 Qi Fan , Hongyu Yuan , Haolin Zuo , Rui Liu , Guanglai Gao

Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comprehension (REC) enables drones to locate objects based on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Zhichao Sun , Yepeng Liu , Zhiling Su , Huachao Zhu , Yuliang Gu , Yuda Zou , Zelong Liu , Gui-Song Xia , Bo Du , Yongchao Xu

Amodal segmentation is a challenging task that aims to predict the complete geometric shape of objects, including their occluded regions. Although existing methods primarily focus on amodal segmentation within the training domain, these…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Bo Zhang , Zhuotao Tian , Xin Tao , Songlin Tang , Jun Yu , Wenjie Pei

We introduce a scalable approach for object pose estimation trained on simulated RGB views of multiple 3D models together. We learn an encoding of object views that does not only describe an implicit orientation of all objects seen during…

Computer Vision and Pattern Recognition · Computer Science 2020-04-06 Martin Sundermeyer , Maximilian Durner , En Yen Puang , Zoltan-Csaba Marton , Narunas Vaskevicius , Kai O. Arras , Rudolph Triebel

Deciphering brain function through non-invasive recordings requires synthesizing complementary high-frequency electromagnetic (EEG/MEG) and low-frequency metabolic (fMRI) signals. However, despite their shared neural origins, extreme…

Neurons and Cognition · Quantitative Biology 2026-02-26 Changli Tang , Shurui Li , Junliang Wang , Qinfan Xiao , Zhonghao Zhai , Lei Bai , Yu Qiao , Bowen Zhou , Wen Wu , Yuanning Li , Chao Zhang

We introduce Multi-Expert Region-based Convolutional Neural Network (ME R-CNN) which is equipped with multiple experts (ME) where each expert is learned to process a certain type of regions of interest (RoIs). This architecture better…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Hyungtae Lee , Sungmin Eum , Heesung Kwon

Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance does not always…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Darryl Hannan , John Cooper , Dylan White , Timothy Doster , Henry Kvinge , Yijing Watkins

Mixture of Vision Encoders (MoVE) has emerged as a powerful approach to enhance the fine-grained visual understanding of multimodal large language models (MLLMs), improving their ability to handle tasks such as complex optical character…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Mozhgan Nasr Azadani , James Riddell , Sean Sedwards , Krzysztof Czarnecki

Latent spaces offer an efficient and effective means of summarizing data while implicitly preserving meta-information through relational encoding. We leverage these meta-embeddings to develop a modality-agnostic, unified encoder. Our method…

Signal Processing · Electrical Eng. & Systems 2025-07-22 Abdullah Ahmed , Jeremy Gummeson
‹ Prev 1 8 9 10 Next ›