中文
相关论文

相关论文: RAMEN: Resolution-Adjustable Multimodal Encoder fo…

200 篇论文

The increasing availability of Earth observation data offers unprecedented opportunities for large-scale environmental monitoring and analysis. However, these datasets are inherently heterogeneous, stemming from diverse sensors,…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Georges Le Bellier , Nicolas Audebert

We address prevailing challenges of the brain-powered research, departing from the observation that the literature hardly recover accurate spatial information and require subject-specific models. To address these challenges, we propose…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Weihao Xia , Raoul de Charette , Cengiz Öztireli , Jing-Hao Xue

The advancement of remote sensing, including satellite systems, facilitates the continuous acquisition of remote sensing imagery globally, introducing novel challenges for achieving open-world tasks. Deployed models need to continuously…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Xiang Xiang , Zhuo Xu , Yao Deng , Qinhao Zhou , Yifan Liang , Ke Chen , Qingfang Zheng , Yaowei Wang , Xilin Chen , Wen Gao

Multimodal variational autoencoders have demonstrated their ability to learn the relationships between different modalities by mapping them into a latent representation. Their design and capacity to perform any-to-any conditional and…

机器学习 · 计算机科学 2025-02-04 Daniel Wesego , Pedram Rooshenas

Masked Autoencoders (MAEs) achieve impressive performance in image classification tasks, yet the internal representations they learn remain less understood. This work started as an attempt to understand the strong downstream classification…

机器学习 · 计算机科学 2026-02-04 Anika Shrivastava , Renu Rameshan , Samar Agnihotri

Computational surface modeling that underlies material recognition has transitioned from reflectance modeling using in-lab controlled radiometric measurements to image-based representations based on internet-mined single-view images…

计算机视觉与模式识别 · 计算机科学 2020-09-24 Jia Xue , Hang Zhang , Ko Nishino , Kristin J. Dana

Existing studies on bundle construction have relied merely on user feedback via bipartite graphs or enhanced item representations using semantic information. These approaches fail to capture elaborate relations hidden in real-world bundle…

Dark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET)…

计算机视觉与模式识别 · 计算机科学 2022-05-09 Ziteng Cui , Guo-Jun Qi , Lin Gu , Shaodi You , Zenghui Zhang , Tatsuya Harada

Recent work by Oberti et al, (Astron. Astrophys., 667, 48, 2022) argued and made a compelling case that classical astronomical adaptive optics (AO) tomography performance can be further enhanced by carefully designing and optically…

天体物理仪器与方法 · 物理学 2025-12-22 Carlos M. Correia , Pierre Jouve , Jesse Cranney , Guido Agapito Cédric Taïssir Heritier

Multimodal learning has been a popular area of research, yet integrating electroencephalogram (EEG) data poses unique challenges due to its inherent variability and limited availability. In this paper, we introduce a novel multimodal…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Kang Yin , Hye-Bin Shin , Dan Li , Seong-Whan Lee

Synthetic Aperture Radar (SAR) and optical imagery provide complementary strengths that constitute the critical foundation for transcending single-modality constraints and facilitating cross-modal collaborative processing and intelligent…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Peihao Wu , Yongxiang Yao , Yi Wan , Wenfei Zhang , Ruipeng Zhao , Jiayuan Li , Yongjun Zhang

Multimodal emotion recognition utilizes complete multimodal information and robust multimodal joint representation to gain high performance. However, the ideal condition of full modality integrity is often not applicable in reality and…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Qi Fan , Hongyu Yuan , Haolin Zuo , Rui Liu , Guanglai Gao

Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comprehension (REC) enables drones to locate objects based on…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zhichao Sun , Yepeng Liu , Zhiling Su , Huachao Zhu , Yuliang Gu , Yuda Zou , Zelong Liu , Gui-Song Xia , Bo Du , Yongchao Xu

Amodal segmentation is a challenging task that aims to predict the complete geometric shape of objects, including their occluded regions. Although existing methods primarily focus on amodal segmentation within the training domain, these…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Bo Zhang , Zhuotao Tian , Xin Tao , Songlin Tang , Jun Yu , Wenjie Pei

We introduce a scalable approach for object pose estimation trained on simulated RGB views of multiple 3D models together. We learn an encoding of object views that does not only describe an implicit orientation of all objects seen during…

计算机视觉与模式识别 · 计算机科学 2020-04-06 Martin Sundermeyer , Maximilian Durner , En Yen Puang , Zoltan-Csaba Marton , Narunas Vaskevicius , Kai O. Arras , Rudolph Triebel

Deciphering brain function through non-invasive recordings requires synthesizing complementary high-frequency electromagnetic (EEG/MEG) and low-frequency metabolic (fMRI) signals. However, despite their shared neural origins, extreme…

神经元与认知 · 定量生物学 2026-02-26 Changli Tang , Shurui Li , Junliang Wang , Qinfan Xiao , Zhonghao Zhai , Lei Bai , Yu Qiao , Bowen Zhou , Wen Wu , Yuanning Li , Chao Zhang

We introduce Multi-Expert Region-based Convolutional Neural Network (ME R-CNN) which is equipped with multiple experts (ME) where each expert is learned to process a certain type of regions of interest (RoIs). This architecture better…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Hyungtae Lee , Sungmin Eum , Heesung Kwon

Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance does not always…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Darryl Hannan , John Cooper , Dylan White , Timothy Doster , Henry Kvinge , Yijing Watkins

Mixture of Vision Encoders (MoVE) has emerged as a powerful approach to enhance the fine-grained visual understanding of multimodal large language models (MLLMs), improving their ability to handle tasks such as complex optical character…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Mozhgan Nasr Azadani , James Riddell , Sean Sedwards , Krzysztof Czarnecki

Latent spaces offer an efficient and effective means of summarizing data while implicitly preserving meta-information through relational encoding. We leverage these meta-embeddings to develop a modality-agnostic, unified encoder. Our method…

信号处理 · 电气工程与系统科学 2025-07-22 Abdullah Ahmed , Jeremy Gummeson
‹ 上一页 1 8 9 10 下一页 ›