English
Related papers

Related papers: Pretrain Where? Investigating How Pretraining Data…

200 papers

The prevailing paradigm for enhancing the reasoning abilities of LLMs revolves around post-training on high-quality, reasoning-intensive data. While emerging literature suggests that reasoning data is increasingly incorporated also during…

Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model.…

Machine Learning · Computer Science 2024-03-12 Hao Chen , Jindong Wang , Ankit Shah , Ran Tao , Hongxin Wei , Xing Xie , Masashi Sugiyama , Bhiksha Raj

Learning meaningful protein representation is important for a variety of biological downstream tasks such as structure-based drug design. Having witnessed the success of protein sequence pretraining, pretraining for structural data which is…

Machine Learning · Computer Science 2023-02-23 Yufei Huang , Lirong Wu , Haitao Lin , Jiangbin Zheng , Ge Wang , Stan Z. Li

Datasets for training object recognition systems are steadily increasing in size. This paper investigates the question of whether existing detectors will continue to improve as data grows, or saturate in performance due to limited model…

Computer Vision and Pattern Recognition · Computer Science 2015-03-06 Xiangxin Zhu , Carl Vondrick , Charless Fowlkes , Deva Ramanan

The deep learning-based analysis of medical images suffers from data scarcity because of high annotation costs and privacy concerns. Researchers in this domain have used transfer learning to avoid overfitting when using complex…

Image and Video Processing · Electrical Eng. & Systems 2022-04-29 Yasar Mehmood , Usama Ijaz Bajwa , Xianfang Sun

Transfer learning is widely used to adapt large pretrained models to new tasks with only a small amount of new data. However, a challenge persists -- the features from the original task often do not fully cover what is needed for unseen…

Machine Learning · Computer Science 2026-02-10 Xingyu Alice Yang , Jianyu Zhang , Léon Bottou

The proliferation of pretrained models, as a result of advancements in pretraining techniques, has led to the emergence of a vast zoo of publicly available models. Effectively utilizing these resources to obtain models with robust…

Machine Learning · Computer Science 2023-06-06 Yimeng Chen , Tianyang Hu , Fengwei Zhou , Zhenguo Li , Zhiming Ma

Significant progress in the development of highly adaptable and reusable Artificial Intelligence (AI) models is expected to have a significant impact on Earth science and remote sensing. Foundation models are pre-trained on large unlabeled…

He et al. (2018) have called into question the utility of pre-training by showing that training from scratch can often yield similar performance to pre-training. We show that although pre-training may not improve performance on traditional…

Machine Learning · Computer Science 2019-10-22 Dan Hendrycks , Kimin Lee , Mantas Mazeika

We empirically demonstrate that a transformer pre-trained on country-scale unlabeled human mobility data learns embeddings capable, through fine-tuning, of developing a deep understanding of the target geography and its corresponding…

Computers and Society · Computer Science 2024-12-13 Alameen Najjar

This paper challenges the recent paradigm in atomic property prediction that links progress to growing dataset sizes and computational resources. We show that pretraining on a carefully selected task-aligned dataset can match or even…

Machine Learning · Computer Science 2026-02-03 Yasir Ghunaim , Hasan Abed Al Kader Hammoud , Bernard Ghanem

Spatial models are used in a variety research areas, such as environmental sciences, epidemiology, or physics. A common phenomenon in many spatial regression models is spatial confounding. This phenomenon takes place when spatially indexed…

Methodology · Statistics 2021-06-08 Isa Marques , Thomas Kneib , Nadja Klein

Earth System Models (ESM) are our main tool for projecting the impacts of climate change. However, running these models at sufficient resolution for local-scale risk-assessments is not computationally feasible. Deep learning-based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Paula Harder , Luca Schmidt , Francis Pelletier , Nicole Ludwig , Matthew Chantry , Christian Lessig , Alex Hernandez-Garcia , David Rolnick

Modeling environmental ecosystems is essential for effective resource management, sustainable development, and understanding complex ecological processes. However, traditional methods frequently struggle with the inherent complexity,…

Machine Learning · Computer Science 2025-03-06 Runlong Yu , Shengyu Chen , Yiqun Xie , Xiaowei Jia

Neural networks have revolutionized many empirical fields, yet their application to financial time series forecasting remains controversial. In this study, we demonstrate that the conventional practice of estimating models locally in…

Econometrics · Economics 2025-02-21 Chen Liu , Minh-Ngoc Tran , Chao Wang , Richard Gerlach , Robert Kohn

In the current landscape of foundation model training, there is a significant reliance on public domain data, which is nearing exhaustion according to recent research. To further scale up, it is crucial to incorporate collaboration among…

Machine Learning · Computer Science 2024-03-08 Wanru Zhao , Yaxin Du , Nicholas Donald Lane , Siheng Chen , Yanfeng Wang

Accurate trajectory prediction is critical for safe autonomous navigation, yet the impact of dataset design on model performance remains understudied. This work systematically examines how feature selection, cross-dataset transfer, and…

Robotics · Computer Science 2025-07-08 Tobias Demmler , Jakob Häringer , Andreas Tamke , Thao Dang , Alexander Hegai , Lars Mikelsons

The data used during training in any given application space is directly tied to the performance of the system once deployed. While there are many other factors that go into producing high performance models within machine learning, there…

Machine Learning · Computer Science 2024-06-17 William H. Clark , Alan J. Michaels

Language model pretraining involves training on extensive corpora, where data quality plays a pivotal role. In this work, we aim to directly estimate the contribution of data during pretraining and select pretraining data in an efficient…

Computation and Language · Computer Science 2025-08-05 Kashun Shum , Yuzhen Huang , Hongjian Zou , Qi Ding , Yixuan Liao , Xiaoxin Chen , Qian Liu , Junxian He

Predictive models are typically trained on historical data to predict future outcomes. While it is commonly assumed that training on more historical data would improve model performance and robustness, data distribution shifts over time may…

Computers and Society · Computer Science 2025-09-05 Chengyuan Yao , Yunxuan Tang , Christopher Brooks , Rene F. Kizilcec , Renzhe Yu