English
Related papers

Related papers: A Complete Decomposition of KL Error using Refined…

200 papers

The estimation problem in a high regression model with structured sparsity is investigated. An algorithm using a two steps block thresholding procedure called GR-LOL is provided. Convergence rates are produced: they depend on simple…

Statistics Theory · Mathematics 2012-07-10 Mathilde Mougeot , Dominique Picard , Karine Tribouley

$L_1$ regularized logistic regression has now become a workhorse of data mining and bioinformatics: it is widely used for many classification problems, particularly ones with many features. However, $L_1$ regularization typically selects…

Machine Learning · Statistics 2015-02-12 Zhe Liu

Bayesian reasoning in linear mixed-effects models (LMMs) is challenging and often requires advanced sampling techniques like Markov chain Monte Carlo (MCMC). A common approach is to write the model in a probabilistic programming language…

Machine Learning · Computer Science 2025-03-25 Jinlin Lai , Justin Domke , Daniel Sheldon

We propose a novel, efficient approach for distributed sparse learning in high-dimensions, where observations are randomly partitioned across machines. Computationally, at each round our method only requires the master machine to solve a…

Machine Learning · Statistics 2016-05-26 Jialei Wang , Mladen Kolar , Nathan Srebro , Tong Zhang

Multiple Choice Question Answering (MCQA) is an important problem with numerous real-world applications, such as medicine, law, and education. The high cost of building MCQA datasets makes few-shot learning pivotal in this domain. While…

Computation and Language · Computer Science 2024-12-31 Patrick Sutanto , Joan Santoso , Esther Irawati Setiawan , Aji Prasetya Wibawa

We present a flexible Bayesian semiparametric mixed model for longitudinal data analysis in the presence of potentially high-dimensional categorical covariates. Building on a novel hidden Markov tensor decomposition technique, our proposed…

Methodology · Statistics 2022-08-05 Giorgio Paulon , Peter Müller , Abhra Sarkar

One of the most important empirical findings in microeconometrics is the pervasiveness of heterogeneity in economic behaviour (cf. Heckman 2001). This paper shows that cumulative distribution functions and quantiles of the nonparametric…

Econometrics · Economics 2020-05-19 Juan Carlos Escanciano

Sensor networks increasingly govern modern infrastructure, yet the data they lose are rarely missing in the uniform-random patterns assumed by standard imputation benchmarks. Loop detectors go offline during calibration, roadside cabinets…

Machine Learning · Computer Science 2026-05-19 Keshu Wu , Sixu Li , Zihao Li , Zhiwen Fan , Xiaopeng Li , Yang Zhou

Clustering is widely used in unsupervised learning to find homogeneous groups of observations within a dataset. However, clustering mixed-type data remains a challenge, as few existing approaches are suited for this task. This study…

Machine Learning · Statistics 2025-11-26 Badih Ghattas , Alvaro Sanchez San-Benito

Molecular profiling data (e.g., gene expression) has been used for clinical risk prediction and biomarker discovery. However, it is necessary to integrate other prior knowledge like biological pathways or gene interaction networks to…

Genomics · Quantitative Biology 2016-09-22 Wenwen Min , Juan Liu , Shihua Zhang

We consider the problem of covariance matrix estimation in the presence of latent variables. Under suitable conditions, it is possible to learn the marginal covariance matrix of the observed variables via a tractable convex program, where…

Machine Learning · Statistics 2011-10-17 Gui-Bo Ye , Yuanfeng Wang , Yifei Chen , Xiaohui Xie

Latent Class Analysis (LCA) is widely used to identify unobserved subgroups in social and behavioural sciences. A long-standing challenge for LCA is the interpretability of the latent classes, due to the high complexity of the estimated…

Methodology · Statistics 2026-05-20 Yuxuan Xu , Lea Kaufmann , Yunxiao Chen , Maria Kateri , Irini Moustaki

Both linear mixed models (LMMs) and sparse regression models are widely used in genetics applications, including, recently, polygenic modeling in genome-wide association studies. These two approaches make very different assumptions, so are…

Quantitative Methods · Quantitative Biology 2012-11-16 Xiang Zhou , Peter Carbonetto , Matthew Stephens

Multi-omics data analysis has the potential to discover hidden molecular interactions, revealing potential regulatory and/or signal transduction pathways for cellular processes of interest when studying life and disease systems. One of…

Quantitative Methods · Quantitative Biology 2022-03-18 Arman Hasanzadeh , Ehsan Hajiramezanali , Nick Duffield , Xiaoning Qian

The proliferation of models for networks raises challenging problems of model selection: the data are sparse and globally dependent, and models are typically high-dimensional and have large numbers of latent variables. Together, these…

Social and Information Networks · Computer Science 2014-06-25 Xiaoran Yan , Cosma Rohilla Shalizi , Jacob E. Jensen , Florent Krzakala , Cristopher Moore , Lenka Zdeborova , Pan Zhang , Yaojia Zhu

Predicting protein functional characteristics from structure remains a central problem in protein science, with broad implications from understanding the mechanisms of disease to designing novel therapeutics. Unfortunately, current machine…

Biological Physics · Physics 2024-09-30 Kevin Borisiak , Gian Marco Visani , Armita Nourmohammad

We theoretically analyze the typical learning performance of $\ell_{1}$-regularized linear regression ($\ell_1$-LinR) for Ising model selection using the replica method from statistical mechanics. For typical random regular graphs in the…

Machine Learning · Computer Science 2022-12-07 Xiangming Meng , Tomoyuki Obuchi , Yoshiyuki Kabashima

It has become common to perform kinetic analysis using approximate Koopman operators that transforms high-dimensional time series of observables into ranked dynamical modes. Key to a practical success of the approach is the identification…

Data Analysis, Statistics and Probability · Physics 2023-10-09 Van A. Ngo , Yen Ting Lin , Danny Perez

The analysis of case-control studies with several subtypes of cases is increasingly common, e.g. in cancer epidemiology. For matched designs, we show that a natural strategy is based on a stratified conditional logistic regression model.…

Methodology · Statistics 2019-01-23 Nadim Ballout , Cedric Garcia , Vivian Viallon

Several approaches to graphically representing context-specific relations among jointly distributed categorical variables have been proposed, along with structure learning algorithms. While existing optimization-based methods have limited…

Machine Learning · Statistics 2024-10-17 Felix Leopoldo Rios , Alex Markham , Liam Solus