中文
相关论文

相关论文: The Residual Information Criterion, Corrected

200 篇论文

The selection of features that are relevant for a prediction or classification problem is an important problem in many domains involving high-dimensional data. Selecting features helps fighting the curse of dimensionality, improving the…

机器学习 · 计算机科学 2009-09-04 Michel Verleysen , Fabrice Rossi , Damien François

A well-known problem in computing some matrix functions iteratively is the lack of a clear, commonly accepted residual notion. An important matrix function for which this is the case is the matrix exponential. Suppose the matrix exponential…

数值分析 · 数学 2015-03-19 Mike A. Botchev

Rating systems are ubiquitous, with applications ranging from product recommendation to teaching evaluations. Confidence intervals for functionals of rating data such as empirical means or quantiles are critical to decision-making in…

统计理论 · 数学 2019-12-10 Robert Nowak , Ervin Tánczos

We introduce a criterion, resilience, which allows properties of a dataset (such as its mean or best low rank approximation) to be robustly computed, even in the presence of a large fraction of arbitrary additional data. Resilience is a…

机器学习 · 计算机科学 2017-11-28 Jacob Steinhardt , Moses Charikar , Gregory Valiant

Experimental designs are tools which can drastically reduce the number of simulations required by time-consuming computer codes. One strategy for selecting the values of the inputs, whose response is to be observed, is to choose these…

统计理论 · 数学 2009-04-17 Astrid Jourdan , Jessica Franco

Modern computational models in supervised machine learning are often highly parameterized universal approximators. As such, the value of the parameters is unimportant, and only the out of sample performance is considered. On the other hand…

统计计算 · 统计学 2021-11-04 Matthew Dixon , Tyler Ward

When data are incomplete, a random vector Y for the data process together with a binary random vector R for the process that causes missing data, are modelled jointly. We review conditions under which R can be ignored for drawing likelihood…

统计方法学 · 统计学 2019-04-01 John C Galati

In a standard classification framework a set of trustworthy learning data are employed to build a decision rule, with the final aim of classifying unlabelled units belonging to the test set. Therefore, unreliable labelled observations,…

应用统计 · 统计学 2019-11-20 Andrea Cappozzo , Francesca Greselin , Thomas Brendan Murphy

For a parametric model of distributions, the closest distribution in the model to the true distribution located outside the model is considered. Measuring the closeness between two distributions with the Kullback-Leibler (K-L) divergence,…

统计理论 · 数学 2025-10-14 Yo Sheena

Counterfactual learning for dealing with missing-not-at-random data (MNAR) is an intriguing topic in the recommendation literature since MNAR data are ubiquitous in modern recommender systems. Missing-at-random (MAR) data, namely randomized…

机器学习 · 计算机科学 2020-10-20 Zifeng Wang , Xi Chen , Rui Wen , Shao-Lun Huang , Ercan E. Kuruoglu , Yefeng Zheng

Probabilistic classifiers are central for making informed decisions under uncertainty. Based on the maximum expected utility principle, optimal decision rules can be derived using the posterior class probabilities and misclassification…

机器学习 · 计算机科学 2025-03-25 Alexandre Perez-Lebel , Gael Varoquaux , Sanmi Koyejo , Matthieu Doutreligne , Marine Le Morvan

In this paper, the sufficient condition in terms of the RIC and ROC for the stable and robust recovery of signals in both noiseless and noisy settings was established via weighted $l_{1}$ minimization when there is partial prior information…

信息论 · 计算机科学 2016-03-15 Wengu Chen , Yaling Li

Model selection is an indispensable part of data analysis dealing very frequently with fitting and prediction purposes. In this paper, we tackle the problem of model selection in a general linear regression where the parameter matrix…

信号处理 · 电气工程与系统科学 2022-09-19 Prakash B. Gohain , Magnus Jansson

Generating free-text rationales is a promising step towards explainable NLP, yet evaluating such rationales remains a challenge. Existing metrics have mostly focused on measuring the association between the rationale and a given label. We…

计算与语言 · 计算机科学 2023-06-05 Hanjie Chen , Faeze Brahman , Xiang Ren , Yangfeng Ji , Yejin Choi , Swabha Swayamdipta

In this paper, we consider the concept of the residual inaccuracy measure and extend it to its weighted version based on extropy. Properties of this measure are studied and the discrimination principle is applied in the class of…

统计方法学 · 统计学 2025-02-03 M. Hashempour , M. R. Kazemi

We view the Information Bottleneck Principle (IBP: Tishby et al., 1999; Schwartz-Ziv and Tishby, 2017) and Predictive Information Bottleneck Principle (PIBP: Still et al., 2007; Alemi, 2019) as special cases of a family of general…

机器学习 · 计算机科学 2019-12-24 Sayandev Mukherjee

There is a large body of evidence that decision makers frequently depart from Bayesian updating. This paper introduces a model, robust maximum likelihood (RML) updating, where deviations from Bayesian updating are due to multiple…

理论经济学 · 经济学 2025-12-17 Elchin Suleymanov

Data analysis based on information from several sources is common in economic and biomedical studies. This setting is often referred to as the data fusion problem, which differs from traditional missing data problems since no complete data…

统计方法学 · 统计学 2022-04-07 Wei Li , Shanshan Luo , Wangli Xu

Semicontinuous outcomes commonly arise in a wide variety of fields, such as insurance claims, healthcare expenditures, rainfall amounts, and alcohol consumption. Regression models, including Tobit, Tweedie, and two-part models, are widely…

统计方法学 · 统计学 2024-03-26 Lu Yang

We consider a new criterion-based approach to model selection in linear regression. Properties of selection criteria based on p-values of a likelihood ratio statistic are studied for families of linear regression models. We prove that such…

统计理论 · 数学 2012-05-21 Piotr Pokarowski , Jan Mielniczuk , Paweł Teisseyre