English
Related papers

Related papers: Jackstraw Inference for AJIVE Data Integration

200 papers

This paper explores useful modifications of the recent development in contrastive learning via novel probabilistic modeling. We derive a particular form of contrastive loss named Joint Contrastive Learning (JCL). JCL implicitly involves the…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Qi Cai , Yu Wang , Yingwei Pan , Ting Yao , Tao Mei

We study the problem of discovering joinable datasets at scale. We approach the problem from a learning perspective relying on profiles. These are succinct representations that capture the underlying characteristics of the schemata and data…

Databases · Computer Science 2023-06-01 Sergi Nadal , Raquel Panadero , Javier Flores , Oscar Romero

Given genetic variations and various phenotypical traits, such as Magnetic Resonance Imaging (MRI) features, we consider two important and related tasks in biomedical research: i)to select genetic and phenotypical markers for disease…

Machine Learning · Computer Science 2013-10-17 Shandian Zhe , Zenglin Xu , Yuan Qi

A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be…

Machine Learning · Computer Science 2019-09-12 Jonas Mueller , Alex Smola

Integrated IPD-AD analysis, which combines individual participant data (IPD) with aggregate data (AD), is increasingly recognized as an effective strategy for generating more reliable and generalizable inferences from heterogeneous studies.…

Methodology · Statistics 2026-03-03 Ming-Yueh Huang , Jing Qin , Chiung-Yu Huang

Given only data generated by a standard confounding graph with unobserved confounder, the Average Treatment Effect (ATE) is not identifiable. To estimate the ATE, a practitioner must then either (a) collect deconfounded data;(b) run a…

Machine Learning · Statistics 2021-03-09 Kyra Gan , Andrew A. Li , Zachary C. Lipton , Sridhar Tayur

We consider the analysis of high dimensional data given in the form of a matrix with columns consisting of observations and rows consisting of features. Often the data is such that the observations do not reside on a regular grid, and the…

Machine Learning · Statistics 2017-08-22 Gal Mishne , Ronen Talmon , Israel Cohen , Ronald R. Coifman , Yuval Kluger

Data clustering reduces the effective sample size from the number of observations towards the number of clusters. For instrumental variable models this reduced effective sample size makes the instruments more likely to be weak, in the sense…

Econometrics · Economics 2025-10-09 Johannes W. Ligtenberg

Clustering is widely used in unsupervised learning to find homogeneous groups of observations within a dataset. However, clustering mixed-type data remains a challenge, as few existing approaches are suited for this task. This study…

Machine Learning · Statistics 2025-11-26 Badih Ghattas , Alvaro Sanchez San-Benito

Though introduced nearly 50 years ago, the infinitesimal jackknife (IJ) remains a popular modern tool for quantifying predictive uncertainty in complex estimation settings. In particular, when supervised learning ensembles are constructed…

Statistics Theory · Mathematics 2021-06-11 Wei Peng , Lucas Mentch , Leonard Stefanski

Understanding the inner workings of complex machine learning models is a long-standing problem and most recent research has focused on local interpretability. To assess the role of individual input features in a global sense, we explore the…

Machine Learning · Computer Science 2020-10-28 Ian Covert , Scott Lundberg , Su-In Lee

Due to the complexity of cancer, clustering algorithms have been used to disentangle the observed heterogeneity and identify cancer subtypes that can be treated specifically. While kernel based clustering approaches allow the use of more…

Machine Learning · Statistics 2018-11-21 Nora K. Speicher , Nico Pfeifer

We develop a concept of weak identification in linear IV models in which the number of instruments can grow at the same rate or slower than the sample size. We propose a jackknifed version of the classical weak identification-robust…

Econometrics · Economics 2021-10-06 Anna Mikusheva , Liyang Sun

Clustering in high-dimensional settings with severe feature noise remains challenging, especially when only a small subset of dimensions is informative and the final number of clusters is not specified in advance. In such regimes, partition…

Machine Learning · Statistics 2026-04-09 Wan Ping Chen

Using the recently developed mathematical theory of tangles, we re-assess the mathematical foundations for applications of the five factor model in personality tests by a new, mathematically rigorous, quantitative method. Our findings…

Neurons and Cognition · Quantitative Biology 2024-12-02 Hanno von Bergen , Reinhard Diestel

Inferring the causal structure of a set of random variables from a finite sample of the joint distribution is an important problem in science. Recently, methods using additive noise models have been suggested to approach the case of…

Machine Learning · Statistics 2012-07-24 Jonas Peters , Dominik Janzing , Bernhard Schölkopf

The learning of predictive models for data-driven decision support has been a prevalent topic in many fields. However, construction of models that would capture interactions among input variables is a challenging task. In this paper, we…

Machine Learning · Computer Science 2019-05-22 Jiapeng Liu , Milosz Kadzinski , Xiuwu Liao , Xiaoxin Mao

We study the implications of including many covariates in a first-step estimate entering a two-step estimation procedure. We find that a first order bias emerges when the number of \textit{included} covariates is "large" relative to the…

Econometrics · Economics 2018-07-27 Matias D. Cattaneo , Michael Jansson , Xinwei Ma

Classical approaches to analyzing dynamical systems, including bifurcation analysis, can provide invaluable insights into underlying structure of a mathematical model, and the spectrum of all possible dynamical behaviors. However, these…

Populations and Evolution · Quantitative Biology 2018-02-16 Irina Kareva

Motivation: Modern biobanks, with unprecedented sample sizes and phenotypic diversity, have become foundational resources for genomic studies, enabling powerful cross-phenotype and population-scale analyses. As studies grow in complexity,…

Applications · Statistics 2026-04-30 Yiran Li , John Whittaker , Sylvia Richardson , Helene Ruffieux
‹ Prev 1 4 5 6 7 8 10 Next ›