中文
相关论文

相关论文: Sparse estimation for case-control studies with mu…

200 篇论文

Given n observations of a p-dimensional random vector, the covariance matrix and its inverse (precision matrix) are needed in a wide range of applications. Sample covariance (e.g. its eigenstructure) can misbehave when p is comparable to…

统计方法学 · 统计学 2008-07-24 Guilherme V. Rocha , Peng Zhao , Bin Yu

We consider estimation and inference in panel data models with additive unobserved individual specific heterogeneity in a high dimensional setting. The setting allows the number of time varying regressors to be larger than the sample size.…

统计方法学 · 统计学 2017-10-05 Alexandre Belloni , Victor Chernozhukov , Christian Hansen , Damian Kozbur

Sparse coding has been popularly used as an effective data representation method in various applications, such as computer vision, medical imaging and bioinformatics, etc. However, the conventional sparse coding algorithms and its manifold…

计算机视觉与模式识别 · 计算机科学 2013-04-04 Jing-Yan Wang

Structured Latent Attribute Models (SLAMs) are a family of discrete latent variable models widely used in education, psychology, and epidemiology to model multivariate categorical data. A SLAM assumes that multiple discrete latent…

统计方法学 · 统计学 2021-07-12 Yuqi Gu , Gongjun Xu

We consider joint estimation of multiple graphical models arising from heterogeneous and high-dimensional observations. Unlike most previous approaches which assume that the cluster structure is given in advance, an appealing feature of our…

机器学习 · 统计学 2018-01-16 Botao Hao , Will Wei Sun , Yufeng Liu , Guang Cheng

Recent studies in the literature have paid much attention to the sparsity in linear classification tasks. One motivation of imposing sparsity assumption on the linear discriminant direction is to rule out the noninformative features, making…

机器学习 · 统计学 2015-01-13 Dong Xia

We introduce a novel class of Bayesian mixtures for normal linear regression models which incorporates a further Gaussian random component for the distribution of the predictor variables. The proposed cluster-weighted model aims to…

统计方法学 · 统计学 2026-05-26 Panagiotis Papastamoulis , Konstantinos Perrakis

In many biomedical problems, data are often heterogeneous, with samples spanning multiple patient subgroups, where different subgroups may have different disease subtypes, stages, or other medical contexts. These subgroups may be related,…

统计方法学 · 统计学 2022-11-30 Zihan Li , Ziye Luo , Yifan Sun

As one of the most popular statistical and machine learning models, logistic regression with regularization has found wide adoption in biomedicine, social sciences, information technology, and so on. These domains often involve data of…

机器学习 · 计算机科学 2016-09-28 Wenfa Li , Hongzhe Liu , Peng Yang , Wei Xie

Regression by composition provides a flexible framework for constructing conditional distributions through sequential group actions. However, when multiple flows act on the same distribution, the model becomes non-identifiable, leading to…

统计方法学 · 统计学 2026-03-30 Safaa K. Kadhem

Statistical model checking delivers quantitative verification results with statistical guarantees by applying Monte Carlo simulation to formal models. It scales to model sizes and model types that are out of reach for exhaustive, analytical…

计算机科学中的逻辑 · 计算机科学 2025-11-18 Pedro R. D'Argenio , Arnd Hartmanns , Patrick Wienhöft , Mark van Wijk

Additive regression provides an extension of linear regression by modeling the signal of a response as a sum of functions of covariates of relatively low complexity. We study penalized estimation in high-dimensional nonparametric additive…

统计理论 · 数学 2017-04-25 Zhiqiang Tan , Cun-Hui Zhang

Controlling false positives (Type I errors) through statistical hypothesis testing is a foundation of modern scientific data analysis. Existing causal structure discovery algorithms either do not provide Type I error control or cannot scale…

统计方法学 · 统计学 2025-12-29 James Leiner , Brian Manzo , Aaditya Ramdas , Wesley Tansey

Here we propose a novel searching scheme for a tuning parameter in high-dimensional penalized regression methods to address variable selection and modeling when sample sizes are limited compared to the data dimensions. Our method is…

定量方法 · 定量生物学 2020-02-11 Tao Jiang , Stephanie J. London , Mi Kyeong Lee , Josyf C. Mychaleckyj , Alison A. Motsinger-Reif

Variable selection for regression models plays a key role in the analysis of biomedical data. However, inference after selection is not covered by classical statistical frequentist theory which assumes a fixed set of covariates in the…

统计方法学 · 统计学 2021-07-21 Michael Kammer , Daniela Dunkler , Stefan Michiels , Georg Heinze

In many randomized trials, outcomes such as essays or open-ended responses must be manually scored as a preliminary step to impact analysis, a process that is costly and limiting. Model-assisted estimation offers a way to combine surrogate…

统计方法学 · 统计学 2026-02-16 Reagan Mozer , Nicole E. Pashley , Luke Miratrix

Logistic regression (LR) is a widely used classification method for modeling binary outcomes in many medical data classification tasks. Research that collects and combines datasets from various data custodians and jurisdictions can…

机器学习 · 计算机科学 2021-05-17 Ali Reza Ghavamipour , Fatih Turkmen , Xiaoqian Jian

We propose a nonparametric method for detecting nonlinear causal relationship within a set of multidimensional discrete time series, by using sparse additive models (SpAMs). We show that, when the input to the SpAM is a $\beta$-mixing time…

机器学习 · 统计学 2018-04-27 Yingxiang Yang , Adams Wei Yu , Zhaoran Wang , Tuo Zhao

High accuracy in cancer prediction is important to improve the quality of the treatment and to improve the rate of survivability of patients. As the data volume is increasing rapidly in the healthcare research, the analytical challenge…

机器学习 · 计算机科学 2014-03-13 J S Saleema , N Bhagawathi , S Monica , P Deepa Shenoy , K R Venugopal , L M Patnaik

Many industry verticals are confronted with small-sized tabular data. In this low-data regime, it is currently unclear whether the best performance can be expected from simple baselines, or more complex machine learning approaches that…

机器学习 · 计算机科学 2024-05-14 Ricardo Knauer , Erik Rodner