English
Related papers

Related papers: Towards Efficient Benchmarking of Foundation Model…

200 papers

The use of high-dimensional features has become a normal practice in many computer vision applications. The large dimension of these features is a limiting factor upon the number of data points which may be effectively stored and processed,…

Computer Vision and Pattern Recognition · Computer Science 2015-06-18 Sakrapee Paisitkriangkrai , Chunhua Shen , Anton van den Hengel

Due to the increase in computational resources and accessibility of data, an increase in large, deep learning models trained on copious amounts of multi-modal data using self-supervised or semi-supervised learning have emerged. These…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Madeline Chantry Schiappa , Shehreen Azad , Sachidanand VS , Yunhao Ge , Ondrej Miksik , Yogesh S. Rawat , Vibhav Vineet

Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operation. Yet, it remains largely a dataset-specific task, requiring comprehensive training data,…

Machine Learning · Computer Science 2026-04-27 Marco Obermeier , Marco Pruckner , Florian Haselbeck , Andreas Zeiselmair

The foundation model has recently garnered significant attention due to its potential to revolutionize the field of visual representation learning in a self-supervised manner. While most foundation models are tailored to effectively process…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Danfeng Hong , Bing Zhang , Xuyang Li , Yuxuan Li , Chenyu Li , Jing Yao , Naoto Yokoya , Hao Li , Pedram Ghamisi , Xiuping Jia , Antonio Plaza , Paolo Gamba , Jon Atli Benediktsson , Jocelyn Chanussot

Image retrieval enables an efficient search through vast amounts of satellite imagery and returns similar images to a query. Deep learning models can identify images across various semantic concepts without the need for annotations. This…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Benedikt Blumenstiel , Viktoria Moor , Romeo Kienzler , Thomas Brunschwiler

Vision foundation models trained on discretely sampled images achieve strong performance on classification benchmarks, yet whether their representations encode the continuous processes underlying their training data remains unclear. This…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Pritika Vig , Ren-Chin Wu , William Lotter

An increasing share of image and video content is analyzed by machines rather than viewed by humans, and therefore it becomes relevant to optimize codecs for such applications where the analysis is performed remotely. Unfortunately,…

Image and Video Processing · Electrical Eng. & Systems 2020-11-13 Lahiru D. Chamain , Fabien Racapé , Jean Bégaint , Akshay Pushparaja , Simon Feltman

Modeling environmental ecosystems is essential for effective resource management, sustainable development, and understanding complex ecological processes. However, traditional methods frequently struggle with the inherent complexity,…

Machine Learning · Computer Science 2025-03-06 Runlong Yu , Shengyu Chen , Yiqun Xie , Xiaowei Jia

Land-cover classification using remote sensing imagery is an important Earth observation task. Recently, land cover classification has benefited from the development of fully connected neural networks for semantic segmentation. The…

Computer Vision and Pattern Recognition · Computer Science 2020-12-09 Xueqing Deng , Yi Zhu , Yuxin Tian , Shawn Newsam

The field of visual representation learning has seen explosive growth in the past years, but its benefits in robotics have been surprisingly limited so far. Prior work uses generic visual representations as a basis to learn (task-specific)…

Robotics · Computer Science 2023-08-16 Jianren Wang , Sudeep Dasari , Mohan Kumar Srirama , Shubham Tulsiani , Abhinav Gupta

We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoders have relied on a variety of pretraining objectives, each…

Function encoders are a recent technique that learn neural network basis functions to form compact, adaptive representations of Hilbert spaces of functions. We show that function encoders provide a principled connection to feature learning…

Machine Learning · Computer Science 2025-09-26 Su Ann Low , Quentin Rommel , Kevin S. Miller , Adam J. Thorpe , Ufuk Topcu

Foundation models, i.e., very large deep learning models, have demonstrated impressive performances in various language and vision tasks that are otherwise difficult to reach using smaller-size models. The major success of GPT-type of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Yiqun Xie , Zhihao Wang , Weiye Chen , Zhili Li , Xiaowei Jia , Yanhua Li , Ruichen Wang , Kangyang Chai , Ruohan Li , Sergii Skakun

Foundation models have excelled in various tasks but are often evaluated on general benchmarks. The adaptation of these models for specific domains, such as remote sensing imagery, remains an underexplored area. In remote sensing, precise…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Ali Mayladan , Hasan Nasrallah , Hasan Moughnieh , Mustafa Shukor , Ali J. Ghandour

Modern Foundation Models (FMs) are typically trained on corpora spanning a wide range of different data modalities, topics and downstream tasks. Utilizing these models can be very computationally expensive and is out of reach for most…

Machine Learning · Computer Science 2025-06-09 Andrey Zhmoginov , Jihwan Lee , Mark Sandler

Visual-to-auditory sensory substitution devices can assist the blind in sensing the visual environment by translating the visual information into a sound pattern. To improve the translation quality, the task performances of the blind are…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Dong Wang , Xuelong Li , Feiping Nie , Qi Wang

Foundation models offer a promising route to transferable remote sensing representations, but many current approaches depend on very large pretraining datasets and fixed sensor configurations, limiting their suitability for ecological and…

Foundation models leverage large-scale pretraining to capture extensive knowledge, demonstrating generalization in a wide range of language tasks. By comparison, vision foundation models (VFMs) often exhibit uneven improvements across…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Shiqi Huang , Yipei Wang , Natasha Thorley , Alexander Ng , Shaheer Saeed , Mark Emberton , Shonit Punwani , Veeru Kasivisvanathan , Dean Barratt , Daniel Alexander , Yipeng Hu

Earth Observation Foundation Models (EOFMs) have exploded in prevalence as tools for processing the massive volumes of remotely sensed and other earth observation data, and for delivering impact on the many essential earth monitoring tasks.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Ryan P. Demilt , Nicholas LaHaye , Karis Tenneson

Recent advances in self-supervision and contrastive learning have brought the performance of foundation models to unprecedented levels in a variety of tasks. Fueled by this progress, these models are becoming the prevailing approach for a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Leo Fillioux , Julio Silva-Rodríguez , Ismail Ben Ayed , Paul-Henry Cournède , Maria Vakalopoulou , Stergios Christodoulidis , Jose Dolz
‹ Prev 1 4 5 6 7 8 10 Next ›