English
Related papers

Related papers: A diagnosis of the primary difference between Euro…

200 papers

Thompson reports a comparison of data from STRmix and TrueAllele. The data he has arises from different inputs to the two software. If the input data are made more similar the outputs become more similar. Thompson argues that the Analytical…

Other Quantitative Biology · Quantitative Biology 2023-06-21 Tim Kalafut , James Curran , Mike Coble , John Buckleton

As Large Language Models (LLMs) achieve breakthroughs in complex reasoning, Codeforces-based Elo ratings have emerged as a prominent metric for evaluating competitive programming capabilities. However, these ratings are often reported…

Software Engineering · Computer Science 2026-02-06 Shenyu Zheng , Ximing Dong , Xiaoshuang Liu , Gustavo Oliva , Chong Chun Yong , Dayi Lin , Boyuan Chen , Shaowei Wang , Ahmed E. Hassan

Electronic Health Records (EHRs) often lack explicit links between medications and diagnoses, making clinical decision-making and research more difficult. Even when links exist, diagnosis lists may be incomplete, especially during early…

Computation and Language · Computer Science 2025-03-31 Dina Albassam , Adam Cross , Chengxiang Zhai

We study the problem of mismatched likelihood ratio test. We analyze the type-\RNum{1} and \RNum{2} error exponents when the actual distributions generating the observation are different from the distributions used in the test. We derive…

Information Theory · Computer Science 2020-01-14 Parham Boroumand , Albert Guillen i Fabregas

Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks -…

Computation and Language · Computer Science 2026-02-13 Eddie Yang , Dashun Wang

Many DNA profiles recovered from crime scene samples are of a quality that does not allow them to be searched against, nor entered into, databases. We propose a method for the comparison of profiles arising from two DNA samples, one or both…

Methodology · Statistics 2017-04-12 K. Ryan , D. Gareth Williams , David J. Balding

Neural networks trained with ERM (empirical risk minimization) sometimes learn unintended decision rules, in particular when their training data is biased, i.e., when training labels are strongly correlated with undesirable features. To…

Computer Vision and Pattern Recognition · Computer Science 2022-11-07 Inwoo Hwang , Sangjun Lee , Yunhyeok Kwak , Seong Joon Oh , Damien Teney , Jin-Hwa Kim , Byoung-Tak Zhang

To make decisions based on a model fit with auto-encoding variational Bayes (AEVB), practitioners often let the variational distribution serve as a surrogate for the posterior distribution. This approach yields biased estimates of the…

Machine Learning · Statistics 2020-10-23 Romain Lopez , Pierre Boyeau , Nir Yosef , Michael I. Jordan , Jeffrey Regier

Many real-world Electronic Health Record (EHR) data contains a large proportion of missing values. Leaving substantial portion of missing information unaddressed usually causes significant bias, which leads to invalid conclusion to be…

Machine Learning · Computer Science 2020-11-04 Lucas J. Liu , Hongwei Zhang , Jianzhong Di , Jin Chen

Both genetic drift and natural selection cause the frequencies of alleles in a population to vary over time. Discriminating between these two evolutionary forces, based on a time series of samples from a population, remains an outstanding…

Populations and Evolution · Quantitative Biology 2013-12-09 Alison Feder , Sergey Kryazhimskiy , Joshua B. Plotkin

For many applications one wishes to decide whether a certain set of numbers originates from an equiprobability distribution or whether they are unequally distributed. Distributions of relative frequencies may deviate significantly from the…

Genomics · Quantitative Biology 2007-05-23 Thorsten Poeschel , Cornelius Froemmel , Christoph Gille

Whole exome sequencing was performed on HLA-matched stem cell donors and transplant recipients to measure sequence variation contributing to minor histocompatibility antigen differences between the two. A large number of nonsynonymous…

Large language models (LLMs) often exhibit tendencies that diverge from human preferences, such as favoring certain writing styles or producing overly verbose outputs. While crucial for improvement, identifying the factors driving these…

Computation and Language · Computer Science 2025-11-18 Juhyun Oh , Eunsu Kim , Jiseon Kim , Wenda Xu , Inha Cha , William Yang Wang , Alice Oh

An interesting phenomenon arises: Empirical Risk Minimization (ERM) sometimes outperforms methods specifically designed for out-of-distribution tasks. This motivates an investigation into the reasons behind such behavior beyond algorithmic…

Machine Learning · Computer Science 2026-01-21 Hong Zheng , Fei Teng

Machine learning (ML) models can underperform on certain population groups due to choices made during model development and bias inherent in the data. We categorize sources of discrimination in the ML pipeline into two classes: aleatoric…

Machine Learning · Computer Science 2024-04-17 Hao Wang , Luxi He , Rui Gao , Flavio P. Calmon

Endmember variability is an important factor for accurately unveiling vital information relating the pure materials and their distribution in hyperspectral images. Recently, the extended linear mixing model (ELMM) has been proposed as a…

Computer Vision and Pattern Recognition · Computer Science 2017-10-24 Tales Imbiriba , Ricardo Augusto Borsoi , José Carlos Moreira Bermudez

The likelihood ratio (LR) measures the relative weight of forensic data regarding two hypotheses. Several levels of uncertainty arise if frequentist methods are chosen for its assessment: the assumed population model only approximates the…

Applications · Statistics 2016-12-26 Giulia Cereda

We study the scaling of classification error rates with respect to the size of the training dataset. In contrast to classical results where rates are minimax optimal for a problem class, this work starts with the empirical observation that,…

Machine Learning · Statistics 2025-06-04 Pengkun Yang , Jingzhao Zhang

We study the feasibility of identifying epistemic uncertainty (reflecting a lack of knowledge), as opposed to aleatoric uncertainty (reflecting entropy in the underlying distribution), in the outputs of large language models (LLMs) over…

Machine Learning · Computer Science 2024-02-28 Gustaf Ahdritz , Tian Qin , Nikhil Vyas , Boaz Barak , Benjamin L. Edelman

A factor model with a break in its factor loadings is observationally equivalent to a model without changes in the loadings but a change in the variance of its factors. This effectively transforms a structural change problem of high…

Econometrics · Economics 2023-12-06 Jushan Bai , Jiangtao Duan , Xu Han
‹ Prev 1 2 3 10 Next ›