中文
相关论文

相关论文: RAMEN: Resolution-Adjustable Multimodal Encoder fo…

200 篇论文

Deep learning has significantly advanced medical imaging analysis, yet variations in image resolution remain an overlooked challenge. Most methods address this by resampling images, leading to either information loss or computational…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Ashay Patel , Michela Antonelli , Sebastien Ourselin , M. Jorge Cardoso

Despite significant advancements in environment perception capabilities for autonomous driving and intelligent robotics, cameras and LiDARs remain notoriously unreliable in low-light conditions and adverse weather, which limits their…

计算机视觉与模式识别 · 计算机科学 2025-01-31 Lei Cheng , Siyang Cao

Spatiotemporal learning is challenging due to the intricate interplay between spatial and temporal dependencies, the high dimensionality of the data, and scalability constraints. These challenges are further amplified in scientific domains,…

机器学习 · 计算机科学 2025-04-17 David Keetae Park , Xihaier Luo , Guang Zhao , Seungjun Lee , Miruna Oprescu , Shinjae Yoo

Currently nearing human-level performance, Visual Question Answering (VQA) is an emerging area in artificial intelligence. Established as a multi-disciplinary field in machine learning, both computer vision and natural language processing…

计算机视觉与模式识别 · 计算机科学 2021-09-07 Bhanuka Manesha Samarasekara Vitharana Gamage , Lim Chern Hong

Earth observation (EO) systems are essential for mapping, catastrophe monitoring, and resource management, but they have trouble processing and sending large amounts of EO data efficiently, especially for specialized applications like…

Satellite Earth-observation (EO) time series in the optical and microwave ranges of the electromagnetic spectrum are often irregular due to orbital patterns and cloud obstruction. Compositing addresses these issues but loses information…

The manifold assumption for high-dimensional data assumes that the data is generated by varying a set of parameters obtained from a low-dimensional latent space. Deep generative models (DGMs) are widely used to learn data representations in…

机器学习 · 计算机科学 2022-07-19 Krithika Iyer , Riddhish Bhalodia , Shireen Elhabian

Deep reinforcement learning agents are often fragile while humans remain adaptive and flexible to varying scenarios. To bridge this gap, we present EDEN, a biologically inspired navigation framework that integrates learned entorhinal-like…

Remote sensing (RS) images from multiple modalities and platforms exhibit diverse details due to differences in sensor characteristics and imaging perspectives. Existing vision-language research in RS largely relies on relatively…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Huiyang Hu , Peijin Wang , Yingchao Feng , Kaiwen Wei , Wenxin Yin , Wenhui Diao , Mengyu Wang , Hanbo Bi , Kaiyue Kang , Tong Ling , Kun Fu , Xian Sun

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when referencing details across multiple input images. In this work,…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Pengcheng Xu , Peng Tang , Donghao Luo , Xiaobin Hu , Weichu Cui , Qingdong He , Zhennan Chen , Jiangning Zhang , Charles Ling , Boyu Wang

Motion mode (M-mode) recording is an essential part of echocardiography to measure cardiac dimension and function. However, the current diagnosis cannot build an automatic scheme, as there are three fundamental obstructs: Firstly, there is…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Ching-Hsun Tseng , Shao-Ju Chien , Po-Shen Wang , Shin-Jye Lee , Wei-Huan Hu , Bin Pu , Xiao-jun Zeng

Recently, referring image segmentation has aroused widespread interest. Previous methods perform the multi-modal fusion between language and vision at the decoding side of the network. And, linguistic feature interacts with visual feature…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Guang Feng , Zhiwei Hu , Lihe Zhang , Huchuan Lu

Real-world geometry and 3D vision tasks are replete with challenging symmetries that defy tractable analytical expression. In this paper, we introduce Neural Isometries, an autoencoder framework which learns to map the observation space to…

计算机视觉与模式识别 · 计算机科学 2024-10-31 Thomas W. Mitchel , Michael Taylor , Vincent Sitzmann

Open-world 3D semantic occupancy prediction aims to generate a voxelized 3D representation from sensor inputs while recognizing both known and unknown objects. Transferring open-vocabulary knowledge from vision-language models (VLMs) offers…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Peizheng Li , Shuxiao Ding , You Zhou , Qingwen Zhang , Onat Inak , Larissa Triess , Niklas Hanselmann , Marius Cordts , Andreas Zell

Quantum machine learning (QML) has gained increasing attention as a potential solution to address the challenges of computation requirements in the future. Earth observation (EO) has entered the era of Big Data, and the computational…

机器学习 · 计算机科学 2026-02-02 Fan Fan , Yilei Shi , Tobias Guggemos , Xiao Xiang Zhu

We present DeepEarth, a self-supervised multi-modal world model with Earth4D, a novel planetary-scale 4D space-time positional encoder. Earth4D extends 3D multi-resolution hash encoding to include time, efficiently scaling across the planet…

Earth observation (EO), aiming at monitoring the state of planet Earth using remote sensing data, is critical for improving our daily lives and living environment. With a growing number of satellites in orbit, an increasing number of…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Zhitong Xiong , Fahong Zhang , Yi Wang , Yilei Shi , Xiao Xiang Zhu

Elevation maps are commonly used to represent the environment of mobile robots and are instrumental for locomotion and navigation tasks. However, pure geometric information is insufficient for many field applications that require appearance…

机器人学 · 计算机科学 2024-10-28 Gian Erni , Jonas Frey , Takahiro Miki , Matias Mattamala , Marco Hutter

Performing simultaneous localization and mapping (SLAM) in low-visibility conditions, such as environments filled with smoke, dust and transparent objets, has long been a challenging task. Sensors like cameras and Light Detection and…

机器人学 · 计算机科学 2024-12-24 Fuhua Jia , Xiaoying Yang , Mengshen Yang , Yang Li , Hang Xu , Adam Rushworth , Salman Ijaz , Heng Yu , Tianxiang Cui

Real-world exposure correction is fundamentally challenged by spatially non-uniform degradations, where diverse exposure errors frequently coexist within a single image. However, existing exposure correction methods are still largely…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Ao Li , Jiawei Sun , Le Dong , Zhenyu Wang , Weisheng Dong