English
Related papers

Related papers: High-dimensional regression over disease subgroups

200 papers

Longitudinal data often involve heterogeneity, sparse signals, and contamination from response outliers or high-leverage observations especially in biomedical science. Existing methods usually address only part of this problem, either…

Methodology · Statistics 2026-02-26 Yuyao Wang , Yu Lu , Tianni Zhang , Mengfei Ran

Deep Learning (DL) can predict biomarkers from cancer histopathology. Several clinically approved applications use this technology. Most approaches, however, predict categorical labels, whereas biomarkers are often continuous measurements.…

Healthcare datasets present many challenges to both machine learning and statistics as their data are typically heterogeneous, censored, high-dimensional and have missing information. Feature selection is often used to identify the…

Machine Learning · Computer Science 2022-07-06 Annette Spooner , Gelareh Mohammadi , Perminder S. Sachdev , Henry Brodaty , Arcot Sowmya

Background. Alzheimer's disease and related dementia (ADRD) are characterized by multiple and progressive anatomo clinical changes. Yet, modeling changes over disease course from cohort data is challenging as the usual timescales are…

We develop estimation for potentially high-dimensional additive structural equation models. A key component of our approach is to decouple order search among the variables from feature or edge selection in a directed acyclic graph encoding…

Methodology · Statistics 2014-12-02 Peter Bühlmann , Jonas Peters , Jan Ernest

The ADReSS Challenge at INTERSPEECH 2020 defines a shared task through which different approaches to the automated recognition of Alzheimer's dementia based on spontaneous speech can be compared. ADReSS provides researchers with a benchmark…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Saturnino Luz , Fasih Haider , Sofia de la Fuente , Davida Fromm , Brian MacWhinney

We propose Causal Interaction Trees for identifying subgroups of participants that have enhanced treatment effects using observational data. We extend the Classification and Regression Tree algorithm by using splitting criteria that focus…

Methodology · Statistics 2021-12-08 Jiabei Yang , Issa J. Dahabreh , Jon A. Steingrimsson

We address a core problem in causal inference: estimating heterogeneous treatment effects using panel data with general treatment patterns. Many existing methods either do not utilize the potential underlying structure in panel data or have…

Machine Learning · Statistics 2024-06-11 Retsef Levi , Elisabeth Paulson , Georgia Perakis , Emily Zhang

Sparse modelling or model selection with categorical data is challenging even for a moderate number of variables, because one parameter is roughly needed to encode one category or level. The Group Lasso is a well known efficient algorithm…

Methodology · Statistics 2022-11-14 Szymon Nowakowski , Piotr Pokarowski , Wojciech Rejchel , Agnieszka Sołtys

Simultaneous variable selection and statistical inference is challenging in high-dimensional data analysis. Most existing post-selection inference methods require explicitly specified regression models, which are often linear, as well as…

Methodology · Statistics 2026-03-19 Shangyuan Ye , Shauna Rakshe , Ye Liang

Hierarchical Bayesian methods enable information sharing across multiple related regression problems. While standard practice is to model regression parameters (effects) as (1) exchangeable across datasets and (2) correlated to differing…

Methodology · Statistics 2021-07-15 Brian L. Trippe , Hilary K. Finucane , Tamara Broderick

Important objectives in cancer research are the prediction of a patient's risk based on molecular measurements such as gene expression data and the identification of new prognostic biomarkers (e.g. genes). In clinical practice, this is…

Applications · Statistics 2020-04-17 Katrin Madjar , Manuela Zucknick , Katja Ickstadt , Jörg Rahnenführer

Challenges with data in the big-data era include (i) the dimension $p$ is often larger than the sample size $n$ (ii) outliers or contaminated points are frequently hidden and more difficult to detect. Challenge (i) renders most conventional…

Machine Learning · Statistics 2023-09-06 Yijun Zuo

Distribution regression has recently attracted much interest as a generic solution to the problem of supervised learning where labels are available at the group level, rather than at the individual level. Current approaches, however, do not…

Machine Learning · Statistics 2021-01-18 Ho Chung Leon Law , Danica J. Sutherland , Dino Sejdinovic , Seth Flaxman

Modeling the time-series of high-dimensional, longitudinal data is important for predicting patient disease progression. However, existing neural network based approaches that learn representations of patient state, while very flexible, are…

Machine Learning · Computer Science 2021-06-21 Zeshan Hussain , Rahul G. Krishnan , David Sontag

Alzheimer's Disease Analysis Model (ADAM) is a multi-agent reasoning large language model (LLM) framework designed to integrate and analyze multimodal data, including microbiome profiles, clinical datasets, and external knowledge bases, to…

Artificial Intelligence · Computer Science 2025-08-22 Ziyuan Huang , Vishaldeep Kaur Sekhon , Roozbeh Sadeghian , Maria L. Vaida , Cynthia Jo , Doyle Ward , Vanni Bucci , John P. Haran

Model-based trees are used to find subgroups in data which differ with respect to model parameters. In some applications it is natural to keep some parameters fixed globally for all observations while asking if and how other parameters vary…

Computation · Statistics 2025-10-07 Heidi Seibold , Torsten Hothorn , Achim Zeileis

Stacking excessive layers in DNN results in highly underdetermined system when training samples are limited, which is very common in medical applications. In this regard, we present a framework capable of deriving an efficient…

Machine Learning · Computer Science 2025-03-04 Seunghun Baek , Injun Choi , Mustafa Dere , Minjeong Kim , Guorong Wu , Won Hwa Kim

Ensemble learning use multiple algorithms to obtain better predictive performance than any single one of its constituent algorithms could. With growing popularity of deep learning, researchers have started to ensemble them for various…

Machine Learning · Computer Science 2019-05-31 Ning An , Huitong Ding , Jiaoyun Yang , Rhoda Au , Ting Fang Alvin Ang

We present a probabilistic programmed deep kernel learning approach to personalized, predictive modeling of neurodegenerative diseases. Our analysis considers a spectrum of neural and symbolic machine learning approaches, which we assess…

Machine Learning · Computer Science 2021-01-13 Alexander Lavin