English
Related papers

Related papers: On Data Thinning for Model Validation in Small Are…

200 papers

Win statistics have gained increasing popularity as primary analysis methods for clinical trials with hierarchical endpoints (HEs) as primary endpoints. However, existing sample size and power calculation approaches in trial design still…

Methodology · Statistics 2026-05-19 Baoshan Zhang , Huiman X. Barnhart , Yuan Wu , Roland A. Matsouaka

Scoring models support decision-making in financial institutions. Their estimation and evaluation are based on the data of previously accepted applicants with known repayment behavior. This creates sampling bias: the available labeled data…

We develop a Bayesian model-based approach to finite population estimation accounting for spatial dependence. Our innovation here is a framework that achieves inference for finite population quantities in spatial process settings. A key…

Applications · Statistics 2019-11-22 Alec M. Chan-Golston , Sudipto Banerjee , Mark S. Handcock

We explore the feasibility and potential of building a ground-truth-free evaluation model to assess the quality of segmentations generated by the Segment Anything Model (SAM) and its variants in medical imaging. This evaluation model…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Ahjol Senbi , Tianyu Huang , Fei Lyu , Qing Li , Yuhui Tao , Wei Shao , Qiang Chen , Chengyan Wang , Shuo Wang , Tao Zhou , Yizhe Zhang

We develop new stochastic gradient methods for efficiently solving sparse linear regression in a partial attribute observation setting, where learners are only allowed to observe a fixed number of actively chosen attributes per example at…

Optimization and Control · Mathematics 2018-12-04 Tomoya Murata , Taiji Suzuki

In this paper, we consider parametric transformed Fay-Herriot models, and clarify conditions on transformations under which the estimator of the transformation is consistent. It is shown that the dual power transformation satisfies the…

Methodology · Statistics 2016-04-07 Shonosuke Sugasawa , Tatsuya Kubokawa

The estimation of the generalization error of classifiers often relies on a validation set. Such a set is hardly available in few-shot learning scenarios, a highly disregarded shortcoming in the field. In these scenarios, it is common to…

Disaggregation regression has become an important tool in spatial disease mapping for making fine-scale predictions of disease risk from aggregated response data. By including high resolution covariate information and modelling the data…

Applications · Statistics 2020-05-08 Rohan Arambepola , Tim C D Lucas , Anita K Nandi , Peter W Gething , Ewan Cameron

National statistical agencies are regularly required to produce estimates about various subpopulations, formed by demographic and/or geographic classifications, based on a limited number of samples. Traditional direct estimates computed…

Methodology · Statistics 2019-10-29 Shuchi Goyal , Gauri Sankar Datta , Abhyuday Mandal

Sparse feature selection is necessary when we fit statistical models, we have access to a large group of features, don't know which are relevant, but assume that most are not. Alternatively, when the number of features is larger than the…

Applications · Statistics 2017-04-04 Emiliano Diaz

Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases in training data. Current methods often rely on costly LLM-based techniques (e.g.…

Artificial Intelligence · Computer Science 2025-12-12 Nick Jiang , Xiaoqing Sun , Lisa Dunlap , Lewis Smith , Neel Nanda

Spatial classification with limited feature observations has been a challenging problem in machine learning. The problem exists in applications where only a subset of sensors are deployed at certain spots or partial responses are collected…

Machine Learning · Computer Science 2020-09-03 Arpan Man Sainju , Wenchong He , Zhe Jiang , Da Yan , Haiquan Chen

Interval-valued data receives much attention due to its wide applications in the fields of finance, econometrics, meteorology and medicine. However, most regression models developed for interval-valued data assume observations are mutually…

Applications · Statistics 2022-10-31 Tingting Huang

Self-supervised landmark estimation is a challenging task that demands the formation of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To tackle this task, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Kejia Yin , Varshanth R. Rao , Ruowei Jiang , Xudong Liu , Parham Aarabi , David B. Lindell

Mislabeled, duplicated, or biased data in real-world scenarios can lead to prolonged training and even hinder model convergence. Traditional solutions prioritizing easy or hard samples lack the flexibility to handle such a variety…

Machine Learning · Computer Science 2023-11-08 Zhijie Deng , Peng Cui , Jun Zhu

Recent changes in housing costs relative to income are likely to affect people's propensity to Housing Affordability Stress (HAS), which is known to have a detrimental effect on a range of health outcomes. The magnitude of these effects may…

Applications · Statistics 2019-12-24 Koen Simons , Rebecca Bentley , Lyle Gurrin

Sparse autoencoders (SAEs) have lately been used to uncover interpretable latent features in large language models. By projecting dense embeddings into a much higher-dimensional and sparse space, learned features become disentangled and…

Machine Learning · Computer Science 2025-07-30 Viktoria Schuster

Recently, new methods for model assessment, based on subsampling and posterior approximations, have been proposed for scaling leave-one-out cross-validation (LOO) to large datasets. Although these methods work well for estimating predictive…

Methodology · Statistics 2020-08-12 Måns Magnusson , Michael Riis Andersen , Johan Jonasson , Aki Vehtari

Sharpness-Aware Minimization (SAM) is a recent training method that relies on worst-case weight perturbations which significantly improves generalization in various settings. We argue that the existing justifications for the success of SAM…

Machine Learning · Computer Science 2022-06-14 Maksym Andriushchenko , Nicolas Flammarion

Feature selection is important step in machine learning since it has shown to improve prediction accuracy while depressing the curse of dimensionality of high dimensional data. The neural networks have experienced tremendous success in…

Machine Learning · Computer Science 2021-07-13 Peter Bugata , Peter Drotar