English
Related papers

Related papers: On the Population Size Estimation from Dual-record…

200 papers

In this work, the goal is to estimate the abundance of an animal population using data coming from capture-recapture surveys. We leverage the prior knowledge about the population's structure to specify a parsimonious finite mixture model…

In order to estimate the population mean in the presence of both non-response and measurement errors that are uncorrelated, the paper presents some novel estimators employing ranked set sampling by utilizing auxiliary information.Up to the…

Methodology · Statistics 2023-11-06 Rajesh Singh , Anamika Kumari

Diversity has been gaining interest in the NLP community in recent years. At the same time, state-of-the-art transformer models such as ModernBERT use very large pre-training datasets, which are driven by size rather than by diversity. This…

Computation and Language · Computer Science 2026-02-26 Louis Estève , Christophe Servan , Thomas Lavergne , Agata Savary

Multiple imputation (MI) has become popular for analyses with missing data in medical research. The standard implementation of MI is based on the assumption of data being missing at random (MAR). However, for missing data generated by…

Methodology · Statistics 2019-01-03 Tra My Pham , James R Carpenter , Tim P Morris , Angela M Wood , Irene Petersen

The determination of sample size in qualitative research has traditionally relied on the subjective and often ambiguous principle of data saturation, which can lead to inconsistencies and threaten methodological rigor. This study introduces…

Machine Learning · Computer Science 2025-12-10 Hasan Tutar , Caner Erden , Ümit Şentürk

Large language models (LLMs) are increasingly being deployed in high-stakes applications like hiring, yet their potential for unfair decision-making remains understudied in generative and retrieval settings. In this work, we examine the…

Computation and Language · Computer Science 2025-09-05 Preethi Seshadri , Hongyu Chen , Sameer Singh , Seraphina Goldfarb-Tarrant

Subsampling algorithms for various parametric regression models with massive data have been extensively investigated in recent years. However, all existing studies on subsampling heavily rely on clean massive data. In practical…

Statistics Theory · Mathematics 2025-06-11 Jiangshan Ju , Mingqiu Wang , Shengli Zhao

Distribution system residential load modeling and analysis for different geographic areas within a utility or an independent system operator territory are critical for enabling small-scale, aggregated distributed energy resources to…

Systems and Control · Electrical Eng. & Systems 2022-08-16 Isaac Bromley-Dulfano , Xiangqi Zhu , Barry Mather

Large language models exhibit societal biases associated with demographic information, including race, gender, and others. Endowing such language models with personalities based on demographic data can enable generating opinions that align…

Artificial Intelligence · Computer Science 2024-02-29 Seungjong Sun , Eungu Lee , Dongyan Nan , Xiangying Zhao , Wonbyung Lee , Bernard J. Jansen , Jang Hyun Kim

Accurate estimation of wildlife density is vital for effective ecological monitoring, conservation, and management. Line transect sampling, a central technique in distance sampling, relies on selecting an appropriate detection function to…

Methodology · Statistics 2025-07-16 Midhat M. Edous , Omar M. Eidous

Population pharmacokinetic (PK) modeling methods can be statistically classified as either parametric or nonparametric (NP). Each classification can be divided into maximum likelihood (ML) or Bayesian (B) approaches. In this paper we…

This research deals with massive multiple hypothesis testing. First regarding multiple tests as an estimation problem under a proper population model, an error measurement called Erroneous Rejection Ratio (ERR) is introduced and related to…

Statistics Theory · Mathematics 2007-06-13 Cheng Cheng

To recognize and mitigate harms from large language models (LLMs), we need to understand the prevalence and nuances of stereotypes in LLM outputs. Toward this end, we present Marked Personas, a prompt-based method to measure stereotypes in…

Computation and Language · Computer Science 2023-05-30 Myra Cheng , Esin Durmus , Dan Jurafsky

Aims: To re-introduce the Heckman model as a valid empirical technique in alcohol studies. Design: To estimate the determinants of problem drinking using a Heckman and a two-part estimation model. Psychological and neuro-scientific studies…

Econometrics · Economics 2023-07-03 Reka Sundaram-Stukel

While machine learning algorithms hold promise for personalised medicine, their clinical adoption remains limited, partly due to biases that can compromise the reliability of predictions. In this paper, we focus on sample selection bias…

Background: The availability of high throughput methods for measurement of mRNA concentrations makes the reliability of conclusions drawn from the data and global quality control of samples and hybridization important issues. We address…

Quantitative Methods · Quantitative Biology 2007-05-23 S. Bilke , T. Breslin , M. Sigvardsson

The paper studies a problem of constructing simultaneous likelihood-based confidence sets. We consider a simultaneous multiplier bootstrap procedure for estimating the quantiles of the joint distribution of the likelihood ratio statistics,…

Statistics Theory · Mathematics 2015-06-19 Mayya Zhilova

Subsampling from a large data set is useful in many supervised learning contexts to provide a global view of the data based on only a fraction of the observations. Diverse (or space-filling) subsampling is an appealing subsampling approach…

Methodology · Statistics 2023-11-27 Boyang Shang , Daniel W. Apley , Sanjay Mehrotra

Measuring customer experience on mobile data is of utmost importance for global mobile operators. The reference signal received power (RSRP) is one of the important indicators for current mobile network management, evaluation and…

Networking and Internet Architecture · Computer Science 2022-07-04 Peizheng Li , Xiaoyang Wang , Robert Piechocki , Shipra Kapoor , Angela Doufexi , Arjun Parekh

Sampling hidden populations is particularly challenging using standard sampling methods mainly because of the lack of a sampling frame. Respondent-driven sampling (RDS) is an alternative methodology that exploits the social contacts between…

‹ Prev 1 8 9 10 Next ›