中文
相关论文

相关论文: Sparse modeling of categorial explanatory variable…

200 篇论文

The paper considers parameter estimation in count data models using penalized likelihood methods. The motivating data consists of multiple independent count variables with a moderate sample size per variable. The data were collected during…

统计方法学 · 统计学 2026-04-15 Minh Thu Bui , Cornelis J. Potgieter , Akihito Kamata

It is known that the common factors in a large panel of data can be consistently estimated by the method of principal components, and principal components can be constructed by iterative least squares regressions. Replacing least squares…

统计方法学 · 统计学 2017-11-16 Jushan Bai , Serena Ng

House price valuation remains challenging due to localized market variations. Existing approaches often rely on black-box machine learning models, which lack interpretability, or simplistic methods like linear regression (LR), which fail to…

机器学习 · 计算机科学 2025-08-06 Paul Gümmer , Julian Rosenberger , Mathias Kraus , Patrick Zschech , Nico Hambauer

We study a panel data model with general heterogeneous effects where slopes are allowed to vary across both individuals and over time. The key dimension reduction assumption we employ is that the heterogeneous slopes can be expressed as…

统计理论 · 数学 2019-09-05 Victor Chernozhukov , Christian Hansen , Yuan Liao , Yinchu Zhu

Penalized logistic regression is extremely useful for binary classification with large number of covariates (higher than the sample size), having several real life applications, including genomic disease classification. However, the…

统计方法学 · 统计学 2023-04-10 Ayanendranath Basu , Abhik Ghosh , María Jaenada , Leandro Pardo

Post-selection inference is a statistical technique for determining salient variables after model or variable selection. Recently, selective inference, a kind of post-selection inference framework, has garnered the attention in the…

统计方法学 · 统计学 2019-06-28 Yuta Umezu , Ichiro Takeuchi

Model-assisted estimation with complex survey data is an important practical problem in survey sampling. When there are many auxiliary variables, selecting significant variables associated with the study variable would be necessary to…

统计方法学 · 统计学 2020-04-01 Shonosuke Sugasawa , Jae Kwang Kim

Dimension reduction and data quantization are two important methods for reducing data complexity. In the paper, we study the methodology of first reducing data dimension by random projection and then quantizing the projections to ternary or…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Weizhi Lu , Mingrui Chen , Kai Guo , Weiyu Li

This paper introduces a simple principle for robust high-dimensional statistical inference via an appropriate shrinkage on the data. This widens the scope of high-dimensional techniques, reducing the moment conditions from sub-exponential…

统计理论 · 数学 2017-05-08 Jianqing Fan , Weichen Wang , Ziwei Zhu

Motivated by the proliferation of observational datasets and the need to integrate non-randomized evidence with randomized controlled trials, causal inference researchers have recently proposed several new methodologies for combining biased…

统计方法学 · 统计学 2023-09-14 Evan T. R. Rosenman , Francesca Dominici , Luke Miratrix

We present a forward sufficient dimension reduction method for categorical or ordinal responses by extending the outer product of gradients and minimum average variance estimator to multinomial generalized linear model. Previous work in…

统计方法学 · 统计学 2023-03-30 Harris Quach , Bing Li

In this paper, we consider Bayesian variable selection problem of linear regression model with global-local shrinkage priors on the regression coefficients. We propose a variable selection procedure that select a variable if the ratio of…

统计方法学 · 统计学 2016-05-26 Xueying Tang , Xiaofan Xu , Malay Ghosh , Prasenjit Ghosh

In a standard classification framework a set of trustworthy learning data are employed to build a decision rule, with the final aim of classifying unlabelled units belonging to the test set. Therefore, unreliable labelled observations,…

应用统计 · 统计学 2019-11-20 Andrea Cappozzo , Francesca Greselin , Thomas Brendan Murphy

We consider the two-group classification problem and propose a kernel classifier based on the optimal scoring framework. Unlike previous approaches, we provide theoretical guarantees on the expected risk consistency of the method. We also…

机器学习 · 统计学 2021-04-01 Alexander F. Lapanowski , Irina Gaynanova

We consider both $\ell _{0}$-penalized and $\ell _{0}$-constrained quantile regression estimators. For the $\ell _{0}$-penalized estimator, we derive an exponential inequality on the tail probability of excess quantile prediction risk and…

统计方法学 · 统计学 2023-03-30 Le-Yu Chen , Sokbae Lee

This paper deals with variable selection in the regression and binary classification frameworks. It proposes an automatic and exhaustive procedure which relies on the use of the CART algorithm and on model selection via penalization. This…

统计理论 · 数学 2011-01-05 Marie Sauvé , Christine Tuleau-Malot

Constrained approaches to maximum likelihood estimation in the context of finite mixtures of normals have been presented in the literature. A fully data-dependent constrained method for maximum likelihood estimation of clusterwise linear…

统计方法学 · 统计学 2016-11-11 Roberto Di Mari , Roberto Rocci , Stefano Antonio Gattone

Multivariate classification methods using explanatory and predictive models are necessary for characterizing subgroups of patients according to their risk profiles. Popular methods include logistic regression and classification trees with…

机器学习 · 计算机科学 2015-11-23 Luca Talenti , Margaux Luck , Anastasia Yartseva , Nicolas Argy , Sandrine Houzé , Cecilia Damon

We investigate the estimation of multivariate extreme models with a discrete spectral measure using spherical clustering techniques. The primary contribution involves devising a method for selecting the order, that is, the number of…

统计方法学 · 统计学 2025-02-20 Shiyuan Deng , He Tang , Shuyang Bai

In statistical machine learning, kernel methods allow to consider infinite dimensional feature spaces with a computational cost that only depends on the number of observations. This is usually done by solving an optimization problem…

最优化与控制 · 数学 2019-01-17 Guillaume Garrigos , Lorenzo Rosasco , Silvia Villa