中文
相关论文

相关论文: What augmentations are sensitive to hyper-paramete…

200 篇论文

Deep neural networks have achieved impressive performance in a wide variety of medical imaging tasks. However, these models often fail on data not used during training, such as data originating from a different medical centre. How to…

图像与视频处理 · 电气工程与系统科学 2022-12-05 Joona Pohjonen , Carolin Stürenberg , Atte Föhr , Reija Randen-Brady , Lassi Luomala , Jouni Lohi , Esa Pitkänen , Antti Rannikko , Tuomas Mirtti

Data augmentation is commonly used to encode invariances in learning methods. However, this process is often performed in an inefficient manner, as artificial examples are created by applying a number of transformations to all points in the…

机器学习 · 计算机科学 2019-03-04 Michael Kuchnik , Virginia Smith

We study the problem of choosing algorithm hyper-parameters in unsupervised domain adaptation, i.e., with labeled data in a source domain and unlabeled data in a target domain, drawn from a different input distribution. We follow the…

This paper aims to understand whether machine learning models should be trained using cost-sensitive surrogates or cost-agnostic ones (e.g., cross-entropy). Analyzing this question through the lens of $\mathcal{H}$-calibration, we find that…

机器学习 · 计算机科学 2025-02-28 Sanket Shah , Milind Tambe , Jessie Finocchiaro

Data augmentation is critical to the empirical success of modern self-supervised representation learning, such as contrastive learning and masked language modeling. However, a theoretical understanding of the exact role of augmentation…

机器学习 · 计算机科学 2024-01-19 Runtian Zhai , Bingbin Liu , Andrej Risteski , Zico Kolter , Pradeep Ravikumar

Data augmentations are important ingredients in the recipe for training robust neural networks, especially in computer vision. A fundamental question is whether neural network features encode data augmentation transformations. To answer…

计算机视觉与模式识别 · 计算机科学 2021-10-29 Eddie Yan , Yanping Huang

In this paper, we investigate the impact of high-dimensional Principal Component (PC) adjustments on inferring the effects of variables on outcomes, with a focus on applications in genetic association studies where PC adjustment is commonly…

统计理论 · 数学 2025-06-30 Sohom Bhattacharya , Rounak Dey , Rajarshi Mukherjee

A main barrier for the deployment of AI radiomic systems in clinical routine is their drop in performance under heterogeneous multicentre acquisition protocols. This work presents a performance-oriented framework for quantifying scan…

人工智能 · 计算机科学 2026-05-15 D. Gil , I. Sanchez , C. Sanchez

Complex statistical models such as scalar-on-image regression often require strong assumptions to overcome the issue of non-identifiability. While in theory it is well understood that model assumptions can strongly influence the results,…

统计方法学 · 统计学 2020-05-04 Clara Happ , Sonja Greven , Volker J. Schmid

We explore unique considerations involved in fitting ML models to data with very high precision, as is often required for science applications. We empirically compare various function approximation methods and study how they scale with…

机器学习 · 计算机科学 2023-02-01 Eric J. Michaud , Ziming Liu , Max Tegmark

For the last two decades, high-dimensional data and methods have proliferated throughout the literature. Yet, the classical technique of linear regression has not lost its usefulness in applications. In fact, many high-dimensional…

Uplift is a particular case of individual treatment effect modeling. Such models deal with cause-and-effect inference for a specific factor, such as a marketing intervention. In practice, these models are built on customer data who…

机器学习 · 统计学 2020-11-03 Belbahri Mouloud , Gandouet Olivier , Kazma Ghaith

Applications of machine learning tools to problems of physical interest are often criticized for producing sensitivity at the expense of transparency. To address this concern, we explore a data planing procedure for identifying combinations…

高能物理 - 唯象学 · 物理学 2018-03-29 Spencer Chang , Timothy Cohen , Bryan Ostdiek

We consider the task of meta-analysis in high-dimensional settings in which the data sources are similar but non-identical. To borrow strength across such heterogeneous datasets, we introduce a global parameter that emphasizes…

统计方法学 · 统计学 2022-07-01 Subha Maity , Yuekai Sun , Moulinath Banerjee

The performance of many machine learning algorithms depends on their hyperparameter settings. The goal of this study is to determine whether it is important to tune a hyperparameter or whether it can be safely set to a default value. We…

机器学习 · 计算机科学 2020-07-16 Hilde J. P. Weerts , Andreas C. Mueller , Joaquin Vanschoren

We improve a known result on the strong consistency of M-estimates of the regression parameters in a linear model for independent and identically distributed random errors under some mild conditions.

统计理论 · 数学 2015-05-28 Xinghui Wang , Shuhe Hu

Locally adapted parameterizations of a model (such as locally weighted regression) are expressive but often suffer from high variance. We describe an approach for reducing the variance, based on the idea of estimating simultaneously a…

机器学习 · 计算机科学 2012-07-03 Doina Precup , Philip Bachman

Uplift modeling is a machine learning technique that aims to model treatment effects heterogeneity. It has been used in business and health sectors to predict the effect of a specific action on a given individual. Despite its advantages,…

机器学习 · 计算机科学 2017-04-20 Atef Shaar , Talel Abdessalem , Olivier Segard

Many problems in computer vision have recently been tackled using models whose predictions cannot be easily interpreted, most commonly deep neural networks. Surrogate explainers are a popular post-hoc interpretability method to further…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Ricardo Kleinlein , Alexander Hepburn , Raúl Santos-Rodríguez , Fernando Fernández-Martínez

The application of deep learning to build accurate predictive models from functional neuroimaging data is often hindered by limited dataset sizes. Though data augmentation can help mitigate such training obstacles, most data augmentation…