中文
相关论文

相关论文: DINO-MX: A Modular & Flexible Framework for Self-S…

200 篇论文

Vision-Language-Action (VLA) models for autonomous driving increasingly adopt generative planners trained with imitation learning followed by reinforcement learning. Diffusion-based planners suffer from modality alignment difficulties, low…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Chenxu Dang , Sining Ang , Yongkang Li , Haochen Tian , Jie Wang , Guang Li , Hangjun Ye , Jie Ma , Long Chen , Yan Wang

Vision foundation models (VFMs) are predominantly developed using data-centric methods. These methods require training on vast amounts of data usually with high-quality labels, which poses a bottleneck for most institutions that lack both…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Jiabo Huang , Chen Chen , Lingjuan Lyu

Vision Foundation Models (VFMs) pretrained on massive datasets exhibit impressive performance on various downstream tasks, especially with limited labeled target data. However, due to their high inference compute cost, these models cannot…

计算机视觉与模式识别 · 计算机科学 2024-07-03 Raviteja Vemulapalli , Hadi Pouransari , Fartash Faghri , Sachin Mehta , Mehrdad Farajtabar , Mohammad Rastegari , Oncel Tuzel

In remote sensing imagery, multi class change detection (MCD) is crucial for fine grained monitoring, yet it has long been constrained by complex scene variations and the scarcity of detailed annotations. To address this, we propose the…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Kai Zheng , Hang-Cheng Dong , Shoulei Liu , Zhenkai Wu , Fupeng Wei , Lei Ding , Wei Zhang

Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Shuhong Zheng , Zhipeng Bao , Ruoyu Zhao , Martial Hebert , Yu-Xiong Wang

Rapid autonomous traversal of unstructured terrain is essential for scenarios such as disaster response, search and rescue, or planetary exploration. As a vehicle navigates at the limit of its capabilities over extreme terrain, its dynamics…

机器人学 · 计算机科学 2024-12-03 Jason Gibson , Anoushka Alavilli , Erica Tevere , Evangelos A. Theodorou , Patrick Spieler

This paper presents a new ambient light normalization framework, DINOLight, that integrates the self-supervised model DINOv2's image understanding capability into the restoration process as a visual prior. Ambient light normalization aims…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Youngjin Oh , Junhyeong Kwon , Nam Ik Cho

Multi-modal Large Language Models (MLLMs) have made significant strides in expanding the capabilities of Large Language Models (LLMs) through the incorporation of visual perception interfaces. Despite the emergence of exciting applications…

计算机视觉与模式识别 · 计算机科学 2024-03-11 Dongsheng Jiang , Yuchen Liu , Songlin Liu , Jin'e Zhao , Hao Zhang , Zhen Gao , Xiaopeng Zhang , Jin Li , Hongkai Xiong

Vision foundation models trained via multi-teacher distillation offer a promising path toward unified visual representations, yet the learning dynamics and data efficiency of such approaches remain underexplored. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Sofian Chaybouti , Sanath Narayan , Yasser Dahou , Phúc H. Lê Khac , Ankit Singh , Ngoc Dung Huynh , Wamiq Reyaz Para , Hilde Kuehne , Hakim Hacid

The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery well without supervised fine-tuning are valuable, but existing solutions fall behind…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Ryan Faulkenberry , Saurabh Prasad

We propose derivative-informed neural operators (DINOs), a general family of neural networks to approximate operators as infinite-dimensional mappings from input function spaces to output function spaces or quantities of interest. After…

数值分析 · 数学 2023-10-18 Thomas O'Leary-Roseberry , Peng Chen , Umberto Villa , Omar Ghattas

Foundation models are vital tools in various Computer Vision applications. They take as input a single RGB image and output a deep feature representation that is useful for various applications. However, in case we have multiple views of…

计算机视觉与模式识别 · 计算机科学 2025-12-18 Leo Segre , Or Hirschorn , Shai Avidan

Foundation vision models are increasingly adopted in medical image analysis. Due to domain shift, these pretrained models misalign with medical image segmentation needs without being fully fine-tuned or lightly adapted. We introduce…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zhuonan Liang , Wei Guo , Jie Gan , Yaxuan Song , Runnan Chen , Hang Chang , Weidong Cai

Numerical simulations play a critical role in design and development of engineering products and processes. Traditional computational methods, such as CFD, can provide accurate predictions but are computationally expensive, particularly for…

This paper presents DINO-SLAM, a DINO-informed design strategy to enhance neural implicit (Neural Radiance Field -- NeRF) and explicit representations (3D Gaussian Splatting -- 3DGS) in SLAM systems through more comprehensive scene…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Ziren Gong , Xiaohan Li , Fabio Tosi , Youmin Zhang , Stefano Mattoccia , Jun Wu , Matteo Poggi

Recent advancements in multimodal vision models have highlighted limitations in late-stage feature fusion and suboptimal query selection for hybrid prompts open-world segmentation, alongside constraints from caption-derived vocabularies. To…

计算机视觉与模式识别 · 计算机科学 2025-08-11 Yuchen Guan , Chong Sun , Canmiao Fu , Zhipeng Huang , Chun Yuan , Chen Li

Vision Foundation Models (VFMs) pretrained on large-scale RGB data have demonstrated remarkable representation quality, yet their applicability to multispectral imaging spanning Near-Infrared (NIR), Short-Wave Infrared (SWIR), and Long-Wave…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Yagiz Nalcakan , Hyeongjin Ju , Incheol Park , Sanghyeop Yeo , Youngwan Jin , Shiho Kim

Multimodal Large Language Models (MLLMs) have achieved remarkable advances by integrating text, image, and audio understanding within a unified architecture. However, existing distributed training frameworks remain fundamentally data-blind:…

分布式、并行与集群计算 · 计算机科学 2026-05-20 Hyeonjun An , Sihyun Kim , Chaerim Lim , Hyunjoon Kim , Rathijit Sen , Sangmin Jung , Hyeonsoo Lee , Dongwook Kim , Takki Yu , Jinkyu Jeong , Youngsok Kim , Kwanghyun Park

Self-supervised monocular depth estimation has gathered notable interest since it can liberate training from dependency on depth annotations. In monocular video training case, recent methods only conduct view synthesis between existing…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Jinfeng Liu , Lingtong Kong , Bo Li , Zerong Wang , Hong Gu , Jinwei Chen

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semantic understanding and…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Jianglong Ye , Naiyan Wang , Xiaolong Wang