English
Related papers

Related papers: Identifying and characterizing extrapolation in mu…

200 papers

We provide a simple formulation of the conditions under which ecological bias should be expected and argue that the bias will affect any method of ecological inference; our claim is supported by formal derivations and several examples where…

Methodology · Statistics 2018-09-25 Antonio Forcina , Davide Pellegrino

In ecological studies niche overlap is often used to quantify species interaction and dynamics. This paper develops a robust, nonparametric statistical framework for quantifying and analyzing multivariate niche overlap. Parametric methods…

Methodology · Statistics 2026-04-08 Jonas Beck , Solomon Harrar

In this present paper, I propose a derivation of unified interpolation and extrapolation function that predicts new values inside and outside the given range by expanding direct Taylor series on the middle point of given data set.…

Numerical Analysis · Mathematics 2020-02-27 Nijat Shukurov

Empirical researchers often estimate spillover effects by fitting linear or non-linear regression models to sampled network data. We show that common sampling schemes bias these estimates, potentially upwards, and derive biased-corrected…

General Economics · Economics 2025-09-23 Kieran Marray

1. Species distribution models and maps from large-scale biodiversity data are necessary for conservation management. One current issue is that biodiversity data are prone to taxonomic misclassifications. Methods to account for these…

Applications · Statistics 2023-05-04 Kwaku Peprah Adjei , Robert B. O'Hara , Wouter Koch , Anders Finstad

We argue that extrapolation to examples outside the training space will often be easier for models that capture global structures, rather than just maximise their local fit to the training data. We show that this is true for two popular…

Computation and Language · Computer Science 2018-05-18 Jeff Mitchell , Pasquale Minervini , Pontus Stenetorp , Sebastian Riedel

When methods of moments are used for identification of power spectral densities, a model is matched to estimated second order statistics such as, e.g., covariance estimates. If the estimates are good there is an infinite family of power…

Optimization and Control · Mathematics 2011-04-12 Per Enqvist

Selective inference methods are developed for group lasso estimators for use with a wide class of distributions and loss functions. The method includes the use of exponential family distributions, as well as quasi-likelihood modeling for…

Methodology · Statistics 2024-03-28 Yiling Huang , Sarah Pirenne , Snigdha Panigrahi , Gerda Claeskens

Calibrating a multiclass predictor, that outputs a distribution over labels, is particularly challenging due to the exponential number of possible prediction values. In this work, we propose a new definition of calibration error that…

Machine Learning · Computer Science 2025-09-30 Konstantina Bairaktari , Huy L. Nguyen

Canonical work handling distribution shifts typically necessitates an entire target distribution that lands inside the training distribution. However, practical scenarios often involve only a handful of target samples, potentially lying…

Machine Learning · Computer Science 2025-01-17 Lingjing Kong , Guangyi Chen , Petar Stojanov , Haoxuan Li , Eric P. Xing , Kun Zhang

Preferential sampling provides a formal modeling specification to capture the effect of bias in a set of sampling locations on inference when a geostatistical model is used to explain observed responses at the sampled locations. In…

Methodology · Statistics 2022-02-21 Shinichiro Shirota , Alan E. Gelfand

In many cases, the values of some model parameters are determined by maximising the likelihood of a set of data points given the parameter values. The presence of outliers in the data and correlations between data points complicate this…

Numerical Analysis · Computer Science 2017-08-28 M. de Jong

Linear regression is a frequently used tool in statistics, however, its validity and interpretability relies on strong model assumptions. While robust estimates of the coefficients' covariance extend the validity of hypothesis tests and…

Methodology · Statistics 2015-04-23 Werner Brannath , Martin Scharpenberg

Multi-label classification consists in classifying an instance into two or more classes simultaneously. It is a very challenging task present in many real-world applications, such as classification of biology, image, video, audio, and text.…

Machine Learning · Computer Science 2020-04-03 Thiago Zafalon Miranda , Diorge Brognara Sardinha , Márcio Porto Basgalupp , Yaochu Jin , Ricardo Cerri

The task of inferring high-level causal variables from low-level observations, commonly referred to as causal representation learning, is fundamentally underconstrained. As such, recent works to address this problem focus on various…

Machine Learning · Statistics 2024-03-26 Simon Bing , Urmi Ninad , Jonas Wahl , Jakob Runge

A critical task in systems biology is the identification of genes that interact to control cellular processes by transcriptional activation of a set of target genes. Many methods have been developed to use statistical correlations in…

Quantitative Methods · Quantitative Biology 2010-11-24 Adam A. Margolin , Kai Wang , Andrea Califano , Ilya Nemenman

It is common in machine learning to estimate a response $y$ given covariate information $x$. However, these predictions alone do not quantify any uncertainty associated with said predictions. One way to overcome this deficiency is with…

Machine Learning · Statistics 2024-06-25 Chancellor Johnstone , Eugene Ndiaye

Probability forecasting is common in the geosciences, the finance sector, and elsewhere. It is sometimes the case that one has multiple probability-forecasts for the same target. How is the information in these multiple forecast systems…

Methodology · Statistics 2016-03-02 Sarah Higgins , Hailiang Du , Leonard A. Smith

We tackle the problem of multi-task learning with copula process. Multivariable prediction in spatial and spatial-temporal processes such as natural resource estimation and pollution monitoring have been typically addressed using techniques…

Machine Learning · Computer Science 2014-06-03 Markus Schneider , Fabio Ramos

The difficulty of multi-class classification generally increases with the number of classes. Using data from a subset of the classes, can we predict how well a classifier will scale with an increased number of classes? Under the assumption…

Machine Learning · Statistics 2016-06-17 Charles Y. Zheng , Rakesh Achanta , Yuval Benjamini
‹ Prev 1 4 5 6 7 8 10 Next ›