English
Related papers

Related papers: Estimation with Binned Data

200 papers

Negative binomial regression is commonly employed to analyze overdispersed count data. With small to moderate sample sizes, the maximum likelihood estimator of the dispersion parameter may be subject to a significant bias, that in turn…

Methodology · Statistics 2020-11-06 Euloge Clovis Kenne Pagui , Alessandra Salvan , Nicola Sartori

Although the methods of bagging and random forests are some of the most widely used prediction methods, relatively little is known about their algorithmic convergence. In particular, there are not many theoretical guarantees for deciding…

Statistics Theory · Mathematics 2019-07-23 Miles E. Lopes

We present a linear regression method for predictions on a small data set making use of a second possibly biased data set that may be much larger. Our method fits linear regressions to the two data sets while penalizing the difference…

Methodology · Statistics 2014-12-19 Aiyou Chen , Art B. Owen , Minghui Shi

Learning ensembles by bagging can substantially improve the generalization performance of low-bias, high-variance estimators, including those evolved by Genetic Programming (GP). To be efficient, modern GP algorithms for evolving (bagging)…

Neural and Evolutionary Computing · Computer Science 2021-02-08 Marco Virgolin

Many areas of the world are without basic information on the socioeconomic well-being of the residing population due to limitations in existing data collection methods. Overhead images obtained remotely, such as from satellite or aircraft,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-14 Ethan Brewer , Giovani Valdrighi , Parikshit Solunke , Joao Rulff , Yurii Piadyk , Zhonghui Lv , Jorge Poco , Claudio Silva

Active learning for regression reduces labeling costs by selecting the most informative samples. Improved Greedy Sampling is a prominent method that balances feature-space diversity and output-space uncertainty using a static,…

Machine Learning · Statistics 2026-03-12 Simon D. Nguyen , Troy Russo , Kentaro Hoffman , Tyler H. McCormick

Pooled logistic regression models are commonly applied in survival analysis. However, the standard implementation can be computationally demanding, which is further exacerbated when using the nonparametric bootstrap for inference. To ease…

Methodology · Statistics 2025-04-21 Paul N Zivich , Stephen R Cole , Bonnie E Shook-Sa , Justin B DeMonte , Jessie K Edwards

Probability distributions defined on the unit interval are widely used in fields ranging from econometrics to reliability studies. Traditional models such as the beta and Kumaraswamy distributions are well-established due to their…

Methodology · Statistics 2026-03-04 Roberto Vila , Helton Saulo , Poliana Matos , Subhankar Dutta

We investigate the forecasting ability of the most commonly used benchmarks in financial economics. We approach the usual caveats of probabilistic forecasts studies -small samples, limited models and non-holistic validations- by performing…

Risk Management · Quantitative Finance 2018-05-08 Ricardo Crisostomo , Lorena Couso

Mixture model-based clustering has become an increasingly popular data analysis technique since its introduction over fifty years ago, and is now commonly utilized within a family setting. Families of mixture models arise when the component…

Methodology · Statistics 2019-11-11 Sanjeena Subedi , Paul D. McNicholas

We analyze and develop a quantitative model describing the evolution of personal income distribution, PID, for males and females in the U.S. between 1930 and 2014. The overall microeconomic model, which we introduced ten years ago,…

General Finance · Quantitative Finance 2015-10-12 Ivan Kitov , Oleg Kitov

Pattern mining is one of the most well-studied subfields in exploratory data analysis. While there is a significant amount of literature on how to discover and rank itemsets efficiently from binary data, there is surprisingly little…

Data Structures and Algorithms · Computer Science 2019-02-05 Nikolaj Tatti

Binary Neural Networks (BiNNs), which employ single-bit precision weights, have emerged as a promising solution to reduce memory usage and power consumption while maintaining competitive performance in large-scale systems. However, training…

Quantum Physics · Physics 2025-11-18 Luca Nepote , Alix Lhéritier , Nicolas Bondoux , Marios Kountouris , Maurizio Filippone

This study downscales the population and gross domestic product (GDP) scenarios given under Shared Socioeconomic Pathways (SSPs) into 0.5-degree grids. Our downscale approach has the following features: (i) it explicitly considers spatial…

Applications · Statistics 2017-04-14 Daisuke Murakami , Yoshiki Yamagata

In machine learning, a bias occurs whenever training sets are not representative for the test data, which results in unreliable models. The most common biases in data are arguably class imbalance and covariate shift. In this work, we aim to…

Machine Learning · Computer Science 2018-04-04 Patrick Glauner , Radu State , Petko Valtchev , Diogo Duarte

Managers, employers, policymakers, and others often seek to understand whether decisions are biased against certain groups. One popular analytic strategy is to estimate disparities after adjusting for observed covariates, typically with a…

Applications · Statistics 2024-01-29 Jongbin Jung , Sam Corbett-Davies , Johann D. Gaebler , Ravi Shroff , Sharad Goel

Modeling large dependent datasets in modern time series analysis is a crucial research area. One effective approach to handle such datasets is to transform the observations into density functions and apply statistical methods for further…

Methodology · Statistics 2025-07-23 Yinzhi Wang , Yingqiu Zhu , Ben-Chang Shia , Lei Qin

In this paper, we introduce a new class of bivariate distributions by compounding the bivariate generalized exponential and power-series distributions. This new class contains some new sub-models such as the bivariate generalized…

Computation · Statistics 2015-08-04 Ali Akbar Jafari , Rasool Roozegar

Change-plane regression identifies subpopulations through an interpretable linear threshold rule, but likelihood-based inference for the hard-threshold boundary is nonregular: objectives are non-smooth, the boundary is weakly identified…

Methodology · Statistics 2026-04-28 Yuki Ohnishi , Fan Li

Small area estimation (SAE) improves estimates for local communities or groups, such as counties, neighborhoods, or demographic subgroups, when data are insufficient for each area. This is important for targeting local resources and…

Methodology · Statistics 2026-01-28 Rayleigh Lei , Yajuan Si
‹ Prev 1 8 9 10 Next ›