English
Related papers

Related papers: Statistically distinguishable rating scale

200 papers

Recommender systems are seen as an effective tool to address information overload, but it is widely known that the presence of various biases makes direct training on large-scale observational data result in sub-optimal prediction…

Information Retrieval · Computer Science 2023-04-19 Haoxuan Li , Yanghao Xiao , Chunyuan Zheng , Peng Wu

In semi-supervised learning, the prevailing understanding suggests that observing additional unlabeled samples improves estimation accuracy for linear parameters only in the case of model misspecification. In this work, we challenge such a…

Methodology · Statistics 2025-09-03 Kai Chen , Yuqian Zhang

Statistical uncertainties complicate engineering design -- confounding regulated design approaches, and degrading the performance of reliability efforts. The simplest means to tackle this uncertainty is double loop simulation; a nested…

Methodology · Statistics 2018-11-02 Zachary del Rosario , Richard W. Fenrich , Gianluca Iaccarino

This manuscript studies statistical properties of linear classifiers obtained through minimization of an unregularized convex risk over a finite sample. Although the results are explicitly finite-dimensional, inputs may be passed through…

Machine Learning · Computer Science 2012-06-15 Matus Telgarsky

Differential privacy is a cryptographically-motivated approach to privacy that has become a very active field of research over the last decade in theoretical computer science and machine learning. In this paradigm one assumes there is a…

Machine Learning · Computer Science 2023-08-02 Marco Avella-Medina

Sharpe ratio (sometimes also referred to as information ratio) is widely used in asset management to compare and benchmark funds and asset managers. It computes the ratio of the (excess) net return over the strategy standard deviation.…

Risk Management · Quantitative Finance 2019-05-22 Eric Benhamou , David Saltiel , Beatrice Guez , Nicolas Paris

Variable selection in the linear regression model takes many apparent faces from both frequentist and Bayesian standpoints. In this paper we introduce a variable selection method referred to as a rescaled spike and slab model. We study the…

Statistics Theory · Mathematics 2007-06-13 Hemant Ishwaran , J. Sunil Rao

In many contemporary applications, large amounts of unlabeled data are readily available while labeled examples are limited. There has been substantial interest in semi-supervised learning (SSL) which aims to leverage unlabeled data to…

Machine Learning · Statistics 2021-09-28 Jessica Gronsbell , Molei Liu , Lu Tian , Tianxi Cai

Measuring the corporate default risk is broadly important in economics and finance. Quantitative methods have been developed to predictively assess future corporate default probabilities. However, as a more difficult yet crucial problem,…

Applications · Statistics 2018-04-26 Miao Yuan , Cheng Yong Tang , Yili Hong , Jian Yang

In recent years, real-world external controls have grown in popularity as a tool to empower randomized placebo-controlled trials, particularly in rare diseases or cases where balanced randomization is unethical or impractical. However, as…

Methodology · Statistics 2024-11-14 Chenyin Gao , Shu Yang , Mingyang Shan , Wenyu Ye , Ilya Lipkovich , Douglas Faries

We present a new approach to assessing the robustness of neural networks based on estimating the proportion of inputs for which a property is violated. Specifically, we estimate the probability of the event that the property is violated…

Machine Learning · Statistics 2019-02-25 Stefan Webb , Tom Rainforth , Yee Whye Teh , M. Pawan Kumar

Semiparametric discrete choice models are widely used in a variety of practical applications. While these models are point identified in the presence of continuous covariates, they can become partially identified when covariates are…

Econometrics · Economics 2024-05-29 Shakeeb Khan , Tatiana Komarova , Denis Nekipelov

We introduce a novel method for sparse regression and variable selection, which is inspired by modern ideas in multiple testing. Imagine we have observations from the linear model y = X beta + z, then we suggest estimating the regression…

Methodology · Statistics 2013-10-30 Malgorzata Bogdan , Ewout van den Berg , Weijie Su , Emmanuel Candes

Binary classification is highly used in credit scoring in the estimation of probability of default. The validation of such predictive models is based both on rank ability, and also on calibration (i.e. how accurately the probabilities…

Econometrics · Economics 2017-10-25 Pedro G. Fonseca , Hugo D. Lopes

The robustness of fault detection algorithms against uncertainty is crucial in the real-world industrial environment. Recently, a new probabilistic design scheme called distributionally robust fault detection (DRFD) has emerged and received…

Optimization and Control · Mathematics 2026-01-16 Yulin Feng , Hailang Jin , Steven X. Ding , Hao Ye , Chao Shang

In this paper we propose and develop a relatively simple and efficient approach for estimating unknown elements of a user-rating matrix in the context of a recommender system (RS). The critical theoretical property of the method is its…

Social and Information Networks · Computer Science 2019-06-04 Jeffrey Uhlmann

Linear and Quadratic Discriminant Analysis are well-known classical methods but can heavily suffer from non-Gaussian distributions and/or contaminated datasets, mainly because of the underlying Gaussian assumption that is not robust. To…

Machine Learning · Statistics 2022-01-11 Pierre Houdouin , Frédéric Pascal , Matthieu Jonckheere , Andrew Wang

We propose a bootstrap-based robust high-confidence level upper bound (Robust H-CLUB) for assessing the risks of large portfolios. The proposed approach exploits rank-based and quantile-based estimators, and can be viewed as a robust…

Statistics Theory · Mathematics 2015-01-13 Jianqing Fan , Fang Han , Han Liu , Byron Vickers

This paper presents a new filter method for unsupervised feature selection. This method is particularly effective on imbalanced multi-class dataset, as in case of clusters of different anomaly types. Existing methods usually involve the…

Machine Learning · Statistics 2023-06-01 Katarina Firdova , Céline Labart , Arthur Martel

The fragility of financial systems was starkly demonstrated in early 2023 through a cascade of major bank failures in the United States, including the second, third, and fourth largest collapses in the US history. The highly interdependent…

Risk Management · Quantitative Finance 2024-11-19 Kamil Fortuna , Janusz Szwabiński