English
Related papers

Related papers: When Data do not Bring Information: A Case Study i…

200 papers

Pre-main sequence (PMS) models provide invaluable tools for the study of star forming regions as they allow to assign masses and ages to young stars. Thus it is of primary importance to test the models against observations of PMS stars with…

Solar and Stellar Astrophysics · Physics 2015-05-30 Mario Gennaro , Pier Giorgio Prada Moroni , Emanuele Tognelli

Statistical models typically capture uncertainties in our knowledge of the corresponding real-world processes, however, it is less common for this uncertainty specification to capture uncertainty surrounding the values of the inputs to the…

Methodology · Statistics 2023-05-10 Samuel E. Jackson , David C. Woods

Zero-shot LLMs are now also used for textual classification tasks, e.g., sentiment and bias detection in a sentence or article. However, their performance can be suboptimal in such data annotation tasks. We introduce a novel technique that…

Computation and Language · Computer Science 2025-11-11 Sina Salimian , Gias Uddin , Shaina Raza , Henry Leung

Histogram-based empirical Bayes methods developed for analyzing data for large numbers of genes, SNPs, or other biological features tend to have large biases when applied to data with a smaller number of features such as genes with…

Methodology · Statistics 2013-10-10 Marta Padilla , David R. Bickel

Two-class classification problems are often characterized by an imbalance between the number of majority and minority datapoints resulting in poor classification of the minority class in particular. Traditional approaches, such as…

Machine Learning · Computer Science 2025-07-11 Karen Medlin , Sven Leyffer , Krishnan Raghavan

Misclassification of binary responses, if ignored, may severely bias the maximum likelihood estimators (MLE) of regression parameters. For such data, a binary regression model incorporating misclassification probabilities is extensively…

Statistics Theory · Mathematics 2020-09-28 Arindam Chatterjee , Tathagata Bandyopadhyay , Sumanta Adhya

Hierarchical parametric models consisting of observable and latent variables are widely used for unsupervised learning tasks. For example, a mixture model is a representative hierarchical model for clustering. From the statistical point of…

Machine Learning · Statistics 2014-01-24 Keisuke Yamazaki

Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal…

Current deep learning classifiers, carry out supervised learning and store class discriminatory information in a set of shared network weights. These weights cannot be easily altered to incrementally learn additional classes, since the…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Penny Johnston , Keiller Nogueira , Kevin Swingler

Several energy management applications rely on accurate photovoltaic generation forecasts. Common metrics like mean absolute error or root-mean-square error, omit error-distribution details needed for stochastic optimization. In addition,…

Machine Learning · Computer Science 2026-03-05 Philipp Danner , Hermann de Meer

Towards understanding the fundamental limits of estimation from data of varied quality, we study the problem of estimating a mean parameter from heteroskedastic Gaussian observations where the variances are unknown and may vary arbitrarily…

Statistics Theory · Mathematics 2026-03-17 Yanjun Han , Abhishek Shetty , Jacob Shkrob

Bayesian model selection provides a formal method of determining the level of support for new parameters in a model. However, if there is not a specific enough underlying physical motivation for the new parameters it can be hard to assign…

Astrophysics · Physics 2009-11-13 Christopher Gordon , Roberto Trotta

We introduce a new, rigorously-formulated Bayesian meta-learning algorithm that learns a probability distribution of model parameter prior for few-shot learning. The proposed algorithm employs a gradient-based variational inference to infer…

Machine Learning · Computer Science 2022-03-21 Cuong Nguyen , Thanh-Toan Do , Gustavo Carneiro

The bottom-up saliency, an early stage of humans' visual attention, can be considered as a binary classification problem between centre and surround classes. Discriminant power of features for the classification is measured as mutual…

Computer Vision and Pattern Recognition · Computer Science 2013-06-07 Anh Cat Le Ngo , Kenneth Li-Minn Ang , Guoping Qiu , Jasmine Kah-Phooi Seng

The $p$-tensor Ising model is a one-parameter discrete exponential family for modeling dependent binary data, where the sufficient statistic is a multi-linear form of degree $p \geq 2$. This is a natural generalization of the matrix Ising…

Statistics Theory · Mathematics 2020-09-01 Somabha Mukherjee , Jaesung Son , Bhaswar B. Bhattacharya

In recent years, Artificial Intelligence techniques have proved to be very successful when applied to problems in physical sciences. Here we apply an unsupervised Machine Learning (ML) algorithm called Principal Component Analysis (PCA) as…

Materials Science · Physics 2021-05-26 T. Tula , G. Möller , J. Quintanilla , S. R. Giblin , A. D. Hillier , E. E. McCabe , S. Ramos , D. S. Barker , S. Gibson

Subsampling is a computationally efficient and scalable method to draw inference in large data settings based on a subset of the data rather than needing to consider the whole dataset. When employing subsampling techniques, a crucial…

Methodology · Statistics 2025-10-08 Amalan Mahendran , Helen Thompson , James M. McGree

We develop two algorithms, based on maximum likelihood (ML) inference, for estimating the parameters of polarized radio sources which emit at a single rotation measure (RM), e.g., pulsars. These algorithms incorporate the flux density…

Instrumentation and Methods for Astrophysics · Physics 2017-10-25 D. H. F. M. Schnitzeler , K. J. Lee

Meta analysis is commonly-used to synthesize multiple results from individual studies. However, its validation is usually threatened by publication bias and between-study heterogeneity, which can be captured by the Copas selection model.…

Methodology · Statistics 2025-07-21 Mengke Li , Yukun Liu , Pengfei Li , Jing Qin

Multilevel linear models allow flexible statistical modelling of complex data with different levels of stratification. Identifying the most appropriate model from the large set of possible candidates is a challenging problem. In the…

Methodology · Statistics 2022-11-15 Tom Edinburgh , Ari Ercole , Stephen J. Eglen