English
Related papers

Related papers: Maven: A Multimodal Foundation Model for Supernova…

200 papers

Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zhekai Chen , Yuqing Wang , Manyuan Zhang , Xihui Liu

Despite their frequent use for change detection, both ConvNets and Vision transformers (ViT) exhibit well-known limitations, namely the former struggle to model long-range dependencies while the latter are computationally inefficient,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Elman Ghazaei , Erchan Aptoula

Multi-view networks are broadly present in real-world applications. In the meantime, network embedding has emerged as an effective representation learning approach for networked data. Therefore, we are motivated to study the problem of…

Social and Information Networks · Computer Science 2019-11-05 Yu Shi , Fangqiu Han , Xinwei He , Xinran He , Carl Yang , Jie Luo , Jiawei Han

Recent advancements in multi-view action recognition have largely relied on Transformer-based models. While effective and adaptable, these models often require substantial computational resources, especially in scenarios with multiple views…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Yuhui Lin , Jiaxuan Lu , Yue Yong , Jiahao Zhang

Pan-sharpening involves integrating information from low-resolution multi-spectral and high-resolution panchromatic images to generate high-resolution multi-spectral counterparts. While recent advancements in the state space model,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Xuanhua He , Ke Cao , Keyu Yan , Rui Li , Chengjun Xie , Jie Zhang , Man Zhou

Mass spectrometry is a widely used method to study molecules and processes in medicine, life sciences, chemistry, catalysis, and industrial product quality control, among many other applications. One of the main features of some mass…

Chemical Physics · Physics 2024-07-02 Daniil A. Boiko , Valentine P. Ananikov

Physical experiments often involve multiple imaging representations, such as X-ray scans and microscopic images. Deep learning models have been widely used for supervised analysis in these experiments. Combining different image…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Nadav Schneider , Muriel Tzdaka , Galit Sturm , Guy Lazovski , Galit Bar , Gilad Oren , Raz Gvishi , Gal Oren

We present OmniSpectra, the first native-resolution foundation model for astronomy spectra. Unlike traditional models, which are limited to fixed-length input sizes or configurations, OmniSpectra handles spectra of any length at their…

Instrumentation and Methods for Astrophysics · Physics 2026-01-23 Md Khairul Islam , Judy Fox

Mesh saliency enhances the adaptability of 3D vision by identifying and emphasizing regions that naturally attract visual attention. To investigate the interaction between geometric structure and texture in shaping visual attention, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Kaiwei Zhang , Dandan Zhu , Xiongkuo Min , Guangtao Zhai

The Platonic Representation Hypothesis claims that recent foundation models are converging to a shared representation space as a function of their downstream task performance, irrespective of the objectives and data modalities used to train…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Laure Ciernik , Lorenz Linhardt , Marco Morik , Jonas Dippel , Simon Kornblith , Lukas Muttenthaler

We present MVMO (Multi-View, Multi-Object dataset): a synthetic dataset of 116,000 scenes containing randomly placed objects of 10 distinct classes and captured from 25 camera locations in the upper hemisphere. MVMO comprises…

Computer Vision and Pattern Recognition · Computer Science 2022-06-01 Aitor Alvarez-Gila , Joost van de Weijer , Yaxing Wang , Estibaliz Garrote

We develop a foundation model using 1.2m high resolution satellite images of the Netherlands. By combining a Convolutional Neural Network and a Vision Transformer, the model captures both low- and high-frequency landscape features, such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Paul Vermeeren , Heysem Kaya

The Phase Extraction Neural Network (PhENN) is a computational architecture, based on deep machine learning, for lens-less quantitative phase retrieval from raw intensity data. PhENN is a deep convolutional neural network trained through…

Image and Video Processing · Electrical Eng. & Systems 2018-11-14 Shuai Li , George Barbastathis

Many vision-related tasks benefit from reasoning over multiple modalities to leverage complementary views of data in an attempt to learn robust embedding spaces. Most deep learning-based methods rely on a late fusion technique whereby…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Austin Reiter , Menglin Jia , Pu Yang , Ser-Nam Lim

Despite the commercial abundance of UAVs, aerial data acquisition remains challenging, and the existing Asia and North America-centric open-source UAV datasets are small-scale or low-resolution and lack diversity in scene contextuality.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Aritra Dutta , Srijan Das , Jacob Nielsen , Rajatsubhra Chakraborty , Mubarak Shah

Large time-domain sky surveys generate extensive multi-year catalogs of light curves in which scientifically valuable transients, such as supernovae (SNe), are vastly outnumbered by artifacts and routine star variability. While supervised…

Instrumentation and Methods for Astrophysics · Physics 2026-03-11 Semenikhin T. A. , Kornilov M. V. , Pruzhinskaya M. V. , Krushinsky V. V. , Malanchev K. L. , Dodin A.

Deep learning architectures are an extremely powerful tool for recognizing and classifying images. However, they require supervised learning and normally work on vectors the size of image pixels and produce the best results when trained on…

Machine Learning · Computer Science 2020-10-20 Ryan Burt , Nina N. Thigpen , Andreas Keil , Jose C. Principe

Existing popular unsupervised embedding learning methods focus on enhancing the instance-level local discrimination of the given unlabeled images by exploring various negative data. However, the existed sample outliers which exhibit large…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Jiahuan Zhou , Yansong Tang , Bing Su , Ying Wu

Diagnosing and treating skin diseases require advanced visual skills across domains and the ability to synthesize information from multiple imaging modalities. While current deep learning models excel at specific tasks like skin cancer…

One of the challenges of studying common neurological disorders is disease heterogeneity including differences in causes, neuroimaging characteristics, comorbidities, or genetic variation. Normative modelling has become a popular method for…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Ana Lawry Aguila , James Chapman , Andre Altmann