中文
相关论文

相关论文: Knowledge-Guided Masked Autoencoder with Linear Sp…

200 篇论文

Photometric consistency loss is one of the representative objective functions commonly used for self-supervised monocular depth estimation. However, this loss often causes unstable depth predictions in textureless or occluded regions due to…

计算机视觉与模式识别 · 计算机科学 2021-11-09 Byeongjun Park , Taekyung Kim , Hyojun Go , Changick Kim

Self-supervised depth estimation has shown its great effectiveness in producing high quality depth maps given only image sequences as input. However, its performance usually drops when estimating on border areas or objects with thin…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Rui Li , Qing Mao , Pei Wang , Xiantuo He , Yu Zhu , Jinqiu Sun , Yanning Zhang

The development of robust and generalisable models for encoding the spatio-temporal dynamics of human brain activity is crucial for advancing neuroscientific discoveries. However, significant individual variation in the organisation of the…

图像与视频处理 · 电气工程与系统科学 2024-06-12 Simon Dahan , Logan Z. J. Williams , Yourong Guo , Daniel Rueckert , Emma C. Robinson

The accurate segmentation of lesions in whole-body PET/CT imaging is es-sential for tumor characterization, treatment planning, and response assess-ment, yet current manual workflows are labor-intensive and prone to inter-observer…

图像与视频处理 · 电气工程与系统科学 2025-09-04 Moona Mazher , Steven A Niederer , Abdul Qayyum

Multimodal representation learning has shown promising improvements on various vision-language tasks. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Zijia Zhao , Longteng Guo , Xingjian He , Shuai Shao , Zehuan Yuan , Jing Liu

Unsupervised spectral unmixing consists of representing each observed pixel as a combination of several pure materials called endmembers with their corresponding abundance fractions. Beyond the linear assumption, various nonlinear unmixing…

计算机视觉与模式识别 · 计算机科学 2023-03-16 Tingting Fang , Fei Zhu , Jie Chen

In self-driving applications, LiDAR data provides accurate information about distances in 3D but lacks the semantic richness of camera data. Therefore, state-of-the-art methods for perception in urban scenes fuse data from both sensor…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Royden Wagner , Marvin Klemp , Carlos Fernandez Lopez

Parameterized mathematical models play a central role in understanding and design of complex information systems. However, they often cannot take into account the intricate interactions innate to such systems. On the contrary, purely…

信号处理 · 电气工程与系统科学 2019-12-02 Shahin Khobahi , Mojtaba Soltanalian

Inspired by the masked language modeling (MLM) in natural language processing tasks, the masked image modeling (MIM) has been recognized as a strong self-supervised pre-training method in computer vision. However, the high random mask ratio…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Zhaowen Li , Yousong Zhu , Zhiyang Chen , Wei Li , Chaoyang Zhao , Rui Zhao , Ming Tang , Jinqiao Wang

The underlying dynamics and patterns of 3D surface meshes deforming over time can be discovered by unsupervised learning, especially autoencoders, which calculate low-dimensional embeddings of the surfaces. To study the deformation patterns…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Sara Hahner , Felix Kerkhoff , Jochen Garcke

This work describes a novel data-driven latent space inference framework built on paired autoencoders to handle observational inconsistencies when solving inverse problems. Our approach uses two autoencoders, one for the parameter space and…

机器学习 · 计算机科学 2026-01-19 Emma Hart , Bas Peters , Julianne Chung , Matthias Chung

We present a new approach for representing and reconstructing multidimensional magnetic resonance imaging (MRI) data. Our method builds on a novel, learned feature-based image representation that disentangles different types of features,…

图像与视频处理 · 电气工程与系统科学 2026-01-01 Ruiyang Zhao , Fan Lam

Accelerated MRI protocols routinely involve a predefined sampling pattern that undersamples the k-space. Finding an optimal pattern can enhance the reconstruction quality, however this optimization is a challenging task. To address this…

图像与视频处理 · 电气工程与系统科学 2024-08-30 Cagan Alkan , Morteza Mardani , Congyu Liao , Zhitao Li , Shreyas S. Vasanawala , John M. Pauly

Depth acquisition, based on active illumination, is essential for autonomous and robotic navigation. LiDARs (Light Detection And Ranging) with mechanical, fixed, sampling templates are commonly used in today's autonomous vehicles. An…

计算机视觉与模式识别 · 计算机科学 2019-08-06 Adam Wolff , Shachar Praisler , Ilya Tcenov , Guy Gilboa

The proliferation of foundation models, pretrained on large-scale unlabeled datasets, has emerged as an effective approach in creating adaptable and reusable architectures that can be leveraged for various downstream tasks using satellite…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Abdul Matin , Tanjim Bin Faruk , Shrideep Pallickara , Sangmi Lee Pallickara

Effective multimodal reasoning depends on the alignment of visual and linguistic representations, yet the mechanisms by which vision-language models (VLMs) achieve this alignment remain poorly understood. Following the LiMBeR framework, we…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Constantin Venhoff , Ashkan Khakzar , Sonia Joseph , Philip Torr , Neel Nanda

Medical vision-and-language pre-training provides a feasible solution to extract effective vision-and-language representations from medical images and texts. However, few studies have been dedicated to this field to facilitate medical…

计算机视觉与模式识别 · 计算机科学 2022-09-16 Zhihong Chen , Yuhao Du , Jinpeng Hu , Yang Liu , Guanbin Li , Xiang Wan , Tsung-Hui Chang

The past year has witnessed a rapid development of masked image modeling (MIM). MIM is mostly built upon the vision transformers, which suggests that self-supervised visual representations can be done by masking input image parts while…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Yunjie Tian , Lingxi Xie , Jiemin Fang , Mengnan Shi , Junran Peng , Xiaopeng Zhang , Jianbin Jiao , Qi Tian , Qixiang Ye

Recent studies have explored using pretrained Vision Foundation Models (VFMs) such as DINO for generative autoencoders, showing strong generative performance. Unfortunately, existing approaches often suffer from limited reconstruction…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Hun Chang , Byunghee Cha , Jong Chul Ye

Autoencoders learn data representations through reconstruction. Robust training is the key factor affecting the quality of the learned representations and, consequently, the accuracy of the application that use them. Previous works…

神经与进化计算 · 计算机科学 2018-07-11 Maisa Doaud , Michael Mayo