English
Related papers

Related papers: Measuring the Instability of Fine-Tuning

200 papers

Forecast evaluation plays a key role in how empirical evidence shapes the development of the discipline. Domain experts are interested in error measures relevant for their decision making needs. Such measures may produce unreliable results.…

Machine Learning · Computer Science 2021-08-10 Hansika Hewamalage , Pablo Montero-Manso , Christoph Bergmeir , Rob J Hyndman

Many modern datasets don't fit neatly into $n \times p$ matrices, but most techniques for measuring statistical stability expect rectangular data. We study methods for stability assessment on non-rectangular data, using statistical learning…

Computation · Statistics 2021-02-23 Kris Sankaran

GPT-3 can perform numerous tasks when provided a natural language prompt that contains a few training examples. We show that this type of few-shot learning can be unstable: the choice of prompt format, training examples, and even the order…

Computation and Language · Computer Science 2021-06-14 Tony Z. Zhao , Eric Wallace , Shi Feng , Dan Klein , Sameer Singh

Incremental learning from non-stationary data poses special challenges to the field of machine learning. Although new algorithms have been developed for this, assessment of results and comparison of behaviors are still open problems, mainly…

Machine Learning · Computer Science 2018-06-19 Alejandro Cervantes , Christian Gagné , Pedro Isasi , Marc Parizeau

Research on bias in machine learning algorithms has generally been concerned with the impact of bias on predictive accuracy. We believe that there are other factors that should also play a role in the evaluation of bias. One such factor is…

Machine Learning · Computer Science 2007-05-23 Peter D. Turney

In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk estimates and its impact on treatment decisions. For overparameterized models, now standard in…

Machine Learning · Computer Science 2026-04-16 Elizabeth W. Miller , Jeffrey D. Blume

Supervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks. In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in…

Computation and Language · Computer Science 2024-10-08 Yiming Ju , Ziyi Ni , Xingrun Xing , Zhixiong Zeng , hanyu Zhao , Siqi Fan , Zheng Zhang

Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet often exhibit poor robustness when exposed to spurious modality interference, such as irrelevant text in vision understanding, or irrelevant…

Machine Learning · Computer Science 2026-01-30 Rui Cai , Bangzheng Li , Xiaofei Wen , Muhao Chen , Zhe Zhao

Owing to their inherently interpretable structure, decision trees are commonly used in applications where interpretability is essential. Recent work has focused on improving various aspects of decision trees, including their predictive…

Machine Learning · Statistics 2023-05-30 Dimitris Bertsimas , Vassilis Digalakis

Multi-Modal Self-Supervised Learning from videos has been shown to improve model's performance on various downstream tasks. However, such Self-Supervised pre-training requires large batch sizes and a large amount of computation resources…

Computer Vision and Pattern Recognition · Computer Science 2021-12-24 Duo Wang , Salah Karout

Predictive Process Monitoring aims to forecast the future progress of process instances using historical event data. As predictive process monitoring is increasingly applied in online settings to enable timely interventions, evaluating the…

Machine Learning · Computer Science 2023-10-16 Suhwan Lee , Marco Comuzzi , Xixi Lu , Hajo A. Reijers

The linear stability of a stratified shear flow for smooth density profiles is studied. This work focuses on the nature of the stability boundaries of flows in which both Kelvin-Helmholtz and Holmboe instabilities are present. For a fixed…

Fluid Dynamics · Physics 2009-11-11 Alexandros Alexakis

We evaluate the robustness of several large language models on multiple datasets. Robustness here refers to the relative insensitivity of the model's answers to meaning-preserving variants of their input. Benchmark datasets are constructed…

Computation and Language · Computer Science 2024-11-05 Samuel Ackerman , Ella Rabinovich , Eitan Farchi , Ateret Anaby-Tavor

Extensively evaluating the capabilities of (large) language models is difficult. Rapid development of state-of-the-art models induce benchmark saturation, while creating more challenging datasets is labor-intensive. Inspired by the recent…

Computation and Language · Computer Science 2025-06-02 Alan Sun

We present an extension to the robust phase estimation protocol, which can identify incorrect results that would otherwise lie outside the expected statistical range. Robust phase estimation is increasingly a method of choice for…

We analyze a numerical instability that occurs in the well-known split-step Fourier method on the background of a soliton. This instability is found to be very sensitive to small changes of the parameters of both the numerical grid and the…

Numerical Analysis · Computer Science 2010-08-31 Taras I. Lakoba

Fine-tuning has become the standard practice for adapting pre-trained models to downstream tasks. However, the impact on model robustness is not well understood. In this work, we characterize the robustness-accuracy trade-off in…

Machine Learning · Computer Science 2025-07-15 Kunyang Li , Jean-Charles Noirot Ferrand , Ryan Sheatsley , Blaine Hoak , Yohan Beugin , Eric Pauley , Patrick McDaniel

Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smaller scales. Although the causes of such instabilities are of…

The sensitivity criterion is widely used in measuring the level of fine-tuning, although many examples show it doesn't work under certain circumstances. We discuss the mathematics behind the fine-tuning problems, explain the mathematical…

High Energy Physics - Phenomenology · Physics 2007-05-23 Su Yan

We discuss recently developed methods that quantify the stability and generalizability of statistical findings under distributional changes. In many practical problems, the data is not drawn i.i.d. from the target population. For example,…

Methodology · Statistics 2023-10-05 Dominik Rothenhäusler , Peter Bühlmann