中文
相关论文

相关论文: Estimation and variable selection in joint mean an…

200 篇论文

Estimating the generalization error (GE) of machine learning models is fundamental, with resampling methods being the most common approach. However, in non-standard settings, particularly those where observations are not independently and…

We present the Mixed Likelihood Gaussian process latent variable model (GP-LVM), capable of modeling data with attributes of different types. The standard formulation of GP-LVM assumes that each observation is drawn from a Gaussian…

机器学习 · 计算机科学 2018-11-20 Samuel Murray , Hedvig Kjellström

We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We…

机器学习 · 统计学 2010-07-16 Lauren A. Hannah , David M. Blei , Warren B. Powell

In this work, we study non-parametric estimation of joint probabilities of a given set of discrete and continuous random variables from their (empirically estimated) 2D marginals, under the assumption that the joint probability could be…

机器学习 · 计算机科学 2022-03-04 Shaan ul Haque , Ajit Rajwade , Karthik S. Gurumoorthy

Variable selection, also known as feature selection in machine learning, plays an important role in modeling high dimensional data and is key to data-driven scientific discoveries. We consider here the problem of detecting influential…

统计方法学 · 统计学 2014-09-24 Bo Jiang , Jun S. Liu

Predictive models trained on imbalanced data tend to produce biased results. This problem is exacerbated when there is not just one output label, but a set of them. This is the case for multilabel learning (MLL) algorithms used to classify…

Additive smooth models, such as Generalized additive models (GAMs) of location, scale, and shape (GAMLSS), are a popular choice for modeling experimental data. However, software available to fit such models is usually not tailored…

统计方法学 · 统计学 2025-06-17 Joshua Krause , Jelmer P. Borst , Jacolien van Rij

In this article, we discuss two specific classes of models - Gaussian Mixture Copula models and Mixture of Factor Analyzers - and the advantages of doing inference with gradient descent using automatic differentiation. Gaussian mixture…

统计计算 · 统计学 2018-12-17 Siva Rajesh Kasa , Vaibhav Rajan

In the last few decades, the study of ordinal data in which the variable of interest is not exactly observed but only known to be in a specific ordinal category has become important. In Psychometrics such variables are analysed under the…

计量经济学 · 经济学 2025-01-22 Bernard M. S. van Praag , J. Peter Hop , William H. Greene

Time series of counts occurring in various applications are often overdispersed, meaning their variance is much larger than the mean. This paper proposes a novel variable selection approach for processing such data. Our approach consists in…

统计方法学 · 统计学 2023-07-04 Marina Gomtsyan

We study uniform consistency in nonparametric mixture models as well as closely related mixture of regression (also known as mixed regression) models, where the regression functions are allowed to be nonparametric and the error…

统计理论 · 数学 2022-12-29 Bryon Aragam , Ruiyi Yang

In this paper, we are concerned with how to select significant variables in semiparametric modeling. Variable selection for semiparametric regression models consists of two components: model selection for nonparametric components and…

统计理论 · 数学 2008-12-18 Runze Li , Hua Liang

Regression models, where the response variable is circular, are common in areas such as biology, geology and meteorology. A typical model assumes that the conditional distribution of the response follows a von-Mises distribution. However,…

统计方法学 · 统计学 2026-01-12 Sphiwe B. Skhosana , Najmeh Nakhaei Rad

In a regression model, prediction is typically performed after model selection. The large variability in the model selection makes the prediction unstable. Thus, it is essential to reduce the variability in model selection and improve…

统计计算 · 统计学 2024-04-11 Wataru Yoshida , Kei Hirose

We investigate the problem of jointly testing two hypotheses and estimating a random parameter based on data that is observed sequentially by sensors in a distributed network. In particular, we assume the data to be drawn from a Gaussian…

信号处理 · 电气工程与系统科学 2020-03-04 Dominik Reinhard , Michael Fauß , Abdelhak M. Zoubir

Variable selection in cluster analysis is important yet challenging. It can be achieved by regularization methods, which realize a trade-off between the clustering accuracy and the number of selected variables by using a lasso-type penalty.…

统计方法学 · 统计学 2016-12-23 Marbac Matthieu , Sedki Mohammed

The Regression Discontinuity Design (RDD) is a quasi-experimental design that estimates the causal effect of a treatment when its assignment is defined by a threshold value for a continuous assignment variable. The RDD assumes that subjects…

应用统计 · 统计学 2020-03-27 Federico Ricciardi , Silvia Liverani , Gianluca Baio

There is an extensive literature on methods for meta-analysis of diagnostic studies, but it mainly focuses on a single test. However, the better understanding of a particular disease has led to the development of multiple tests. A…

统计方法学 · 统计学 2020-10-19 Aristidis K. Nikoloulopoulos

The composite likelihood (CL) is amongst the computational methods used for estimation of the generalized linear mixed model (GLMM) in the context of bivariate meta-analysis of diagnostic test accuracy studies. Its advantage is that the…

统计方法学 · 统计学 2018-07-12 Aristidis K. Nikoloulopoulos

Kernel techniques are among the most popular and flexible approaches in data science allowing to represent probability measures without loss of information under mild conditions. The resulting mapping called mean embedding gives rise to a…

机器学习 · 统计学 2024-11-27 Linda Chamakh , Zoltan Szabo