中文
相关论文

相关论文: Data-Driven Subgroup Identification for Linear Reg…

200 篇论文

A representative model in integrative analysis of two high-dimensional correlated datasets is to decompose each data matrix into a low-rank common matrix generated by latent factors shared across datasets, a low-rank distinctive matrix…

机器学习 · 统计学 2022-04-06 Hai Shu , Zhe Qu

In multicenter biomedical research, integrating data from multiple decentralized sites provides more robust and generalizable findings due to its larger sample size and the ability to account for the between-site heterogeneity. However,…

统计方法学 · 统计学 2025-12-29 Xiaokang Liu , Yuchen Yang , Yifei Sun , Jiang Bian , Yanyuan Ma , Raymond J. Carroll , Yong Chen

Deep neural networks have shown the ability to extract universal feature representations from data such as images and text that have been useful for a variety of learning tasks. However, the fruits of representation learning have yet to be…

机器学习 · 计算机科学 2023-03-28 Liam Collins , Hamed Hassani , Aryan Mokhtari , Sanjay Shakkottai

Machine learning models can fail on subgroups that are underrepresented during training. While techniques such as dataset balancing can improve performance on underperforming groups, they require access to training group annotations and can…

机器学习 · 计算机科学 2024-06-25 Saachi Jain , Kimia Hamidieh , Kristian Georgiev , Andrew Ilyas , Marzyeh Ghassemi , Aleksander Madry

Important objectives in cancer research are the prediction of a patient's risk based on molecular measurements such as gene expression data and the identification of new prognostic biomarkers (e.g. genes). In clinical practice, this is…

应用统计 · 统计学 2020-04-17 Katrin Madjar , Manuela Zucknick , Katja Ickstadt , Jörg Rahnenführer

This paper introduces a new data analysis method for big data using a newly defined regression model named multiple model linear regression(MMLR), which separates input datasets into subsets and construct local linear regression models of…

机器学习 · 计算机科学 2023-08-25 Bohan Lyu , Jianzhong Li

Modern data sets, such as those in healthcare and e-commerce, are often derived from many individuals or systems but have insufficient data from each source alone to separately estimate individual, often high-dimensional, model parameters.…

机器学习 · 计算机科学 2024-11-14 Maryann Rui , Thibaut Horel , Munther Dahleh

Many data sets contain an inherent multilevel structure, for example, because of repeated measurements of the same observational units. Taking this structure into account is critical for the accuracy and calibration of any statistical…

统计方法学 · 统计学 2020-05-07 Topi Paananen , Alejandro Catalina , Paul-Christian Bürkner , Aki Vehtari

A main challenge of data-driven sciences is how to make maximal use of the progressively expanding databases of experimental datasets in order to keep research cumulative. We introduce the idea of a modeling-based dataset retrieval engine…

定量方法 · 定量生物学 2015-06-19 Ali Faisal , Jaakko Peltonen , Elisabeth Georgii , Johan Rung , Samuel Kaski

In this manuscript, we investigate the concept of the mean response for a treatment group mean as well as its estimation and prediction for generalized linear models with a subject-wise random effect. Generalized linear models are commonly…

应用统计 · 统计学 2019-11-05 Jiexin Duan , Michael Levine , Junxiang Luo , Yongming Qu

One important problem in microbiome analysis is to identify the bacterial taxa that are associated with a response, where the microbiome data are summarized as the composition of the bacterial taxa at different taxonomic levels. This paper…

应用统计 · 统计学 2016-03-04 Pixu Shi , Anru Zhang , Hongzhe Li

Research in modern data-driven dynamical systems is typically focused on the three key challenges of high dimensionality, unknown dynamics, and nonlinearity. The dynamic mode decomposition (DMD) has emerged as a cornerstone for modeling…

流体动力学 · 物理学 2022-04-27 Peter J. Baddoo , Benjamin Herrmann , Beverley J. McKeon , Steven L. Brunton

Randomization, as a key technique in clinical trials, can eliminate sources of bias and produce comparable treatment groups. In randomized experiments, the treatment effect is a parameter of general interest. Researchers have explored the…

统计方法学 · 统计学 2023-12-05 Fuyi Tu , Wei Ma , Hanzhong Liu

Circuit discovery aims to explain how language models (LMs) implement a specific task by localizing and interpreting a circuit, a computational subgraph responsible for the LM's behavior. Existing circuit discovery methods are…

人工智能 · 计算机科学 2026-05-12 Daking Rai , Mor Geva , Ziyu Yao

In this paper we explore different regression models based on Clusterwise Linear Regression (CLR). CLR aims to find the partition of the data into $k$ clusters, such that linear regressions fitted to each of the clusters minimize overall…

机器学习 · 计算机科学 2018-05-01 Igor Gitman , Jieshi Chen , Eric Lei , Artur Dubrawski

We propose generalized additive partial linear models for complex data which allow one to capture nonlinear patterns of some covariates, in the presence of linear components. The proposed method improves estimation efficiency and increases…

统计理论 · 数学 2014-05-26 Li Wang , Lan Xue , Annie Qu , Hua Liang

We describe a data-driven discovery method that leverages Simpson's paradox to uncover interesting patterns in behavioral data. Our method systematically disaggregates data to identify subgroups within a population whose behavior deviates…

计算机与社会 · 计算机科学 2018-05-09 Nazanin Alipourfard , Peter G. Fennell , Kristina Lerman

In multivariate time series systems, key insights can be obtained by discovering lead-lag relationships inherent in the data, which refer to the dependence between two time series shifted in time relative to one another, and which can be…

机器学习 · 统计学 2023-09-20 Yichi Zhang , Mihai Cucuringu , Alexander Y. Shestopaloff , Stefan Zohren

Symbolic regression is a machine learning technique that can learn the governing formulas of data and thus has the potential to transform scientific discovery. However, symbolic regression is still limited in the complexity and…

机器学习 · 计算机科学 2023-05-30 Michael Zhang , Samuel Kim , Peter Y. Lu , Marin Soljačić

We address the problem of data-driven pattern identification and outlier detection in time series. To this end, we use singular value decomposition (SVD) which is a well-known technique to compute a low-rank approximation for an arbitrary…

统计方法学 · 统计学 2019-03-12 Abdolrahman Khoshrou , Eric J. Pauwels