English
Related papers

Related papers: A Common-Factor Approach for Multivariate Data Cle…

200 papers

The emergent dynamics of complex systems often arise from the internal dynamical interactions among different elements and hence is to be modeled using multiple variables that represent the different dynamical processes. When such systems…

Chaotic Dynamics · Physics 2024-11-05 Shivam Kumar , R. Misra , G. Ambika

Economists are blessed with a wealth of data for analysis, but more often than not, values in some entries of the data matrix are missing. Various methods have been proposed to handle missing observations in a few variables. We exploit the…

Econometrics · Economics 2022-02-02 Ercument Cahan , Jushan Bai , Serena Ng

Quantitative evaluations of differences and/or similarities between data samples define and shape optimisation problems associated with learning data distributions. Current methods to compare data often suffer from limitations in capturing…

Machine Learning · Computer Science 2024-01-23 Deborah Pelacani Cruz , George Strong , Oscar Bates , Carlos Cueto , Jiashun Yao , Lluis Guasch

Data integration is the problem of combining multiple data groups (studies, cohorts) and/or multiple data views (variables, features). This task is becoming increasingly important in many disciplines due to the prevalence of large and…

Methodology · Statistics 2019-11-13 Jonatan Kallus , Patrik Johansson , Sven Nelander , Rebecka Jörnsten

Machine learning (ML) is transforming modeling and control in the physical, engineering, and biological sciences. However, rapid development has outpaced the creation of standardized, objective benchmarks - leading to weak baselines,…

In myriad statistical applications, data are collected from related but heterogeneous sources. These sources share some commonalities while containing idiosyncratic characteristics. One of the most fundamental challenges in such scenarios…

Methodology · Statistics 2024-03-29 Naichen Shi , Raed Al Kontar , Salar Fattahi

Recent diagnostic datasets on compositional generalization, such as SCAN (Lake and Baroni, 2018) and COGS (Kim and Linzen, 2020), expose severe problems in models trained from scratch on these datasets. However, in contrast to this poor…

Computation and Language · Computer Science 2023-11-09 Xiang Zhou , Yichen Jiang , Mohit Bansal

Researchers often have datasets measuring features $x_{ij}$ of samples, such as test scores of students. In factor analysis and PCA, these features are thought to be influenced by unobserved factors, such as skills. Can we determine how…

Statistics Theory · Mathematics 2019-09-16 Edgar Dobriban

The emergence of large-scale pre-trained vision foundation models has greatly advanced the medical imaging field through the pre-training and fine-tuning paradigm. However, selecting appropriate medical data for downstream fine-tuning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Anyang Ji , Qingbo Kang , Wei Xu , Changfan Wang , Kang Li , Qicheng Lao

The FAIR principles for scientific data (Findable, Accessible, Interoperable, Reusable) are also relevant to other digital objects such as research software and scientific workflows that operate on scientific data. The FAIR principles can…

Digital Libraries · Computer Science 2022-12-16 Sean R. Wilkinson , Greg Eisenhauer , Anuj J. Kapadia , Kathryn Knight , Jeremy Logan , Patrick Widener , Matthew Wolf

Data cleaning is a crucial yet challenging task in data analysis, often requiring significant manual effort. To automate data cleaning, previous systems have relied on statistical rules derived from erroneous data, resulting in low accuracy…

Databases · Computer Science 2024-10-22 Shuo Zhang , Zezhou Huang , Eugene Wu

We consider a general statistical estimation problem involving a finite-dimensional target parameter vector. Beyond an internal data set drawn from the population distribution, external information, such as additional individual data or…

Methodology · Statistics 2025-07-31 Guorong Dai , Lingxuan Shao , Jinbo Chen

Despite the progress made in deepfake detection research, recent studies have shown that biases in the training data for these detectors can result in varying levels of performance across different demographic groups, such as race and…

Machine Learning · Computer Science 2025-01-03 Uzoamaka Ezeakunne , Chrisantus Eze , Xiuwen Liu

Conjoint analysis is a popular experimental design used to measure multidimensional preferences. Researchers examine how varying a factor of interest, while controlling for other relevant factors, influences decision-making. Currently,…

Methodology · Statistics 2024-11-20 Dae Woong Ham , Kosuke Imai , Lucas Janson

We present a new approach to factor rotation for functional data. This is achieved by rotating the functional principal components toward a predefined space of periodic functions designed to decompose the total variation into components…

Applications · Statistics 2012-07-02 Chong Liu , Surajit Ray , Giles Hooker , Mark Friedl

Often in surveys, key items are subject to measurement errors. Given just the data, it can be difficult to determine the distribution of this error process, and hence to obtain accurate inferences that involve the error-prone variables. In…

Methodology · Statistics 2016-10-04 Tracy Schifeling , Jerome P. Reiter , Maria DeYoreo

This article focuses on covariance estimation for multi-study data. Popular approaches employ factor-analytic terms with shared and study-specific loadings that decompose the variance into (i) a shared low-rank component, (ii)…

Methodology · Statistics 2026-01-26 Lorenzo Mauri , Niccolò Anceschi , David B. Dunson

An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable…

Many of the traditional recommendation algorithms are designed based on the fundamental idea of mining or learning correlative patterns from data to estimate the user-item correlative preference. However, pure correlative learning may lead…

Information Retrieval · Computer Science 2023-08-15 Shuyuan Xu , Yingqiang Ge , Yunqi Li , Zuohui Fu , Xu Chen , Yongfeng Zhang

We study the optimization problem of selecting numerical quantities to clean in order to fact-check claims based on such data. Oftentimes, such claims are technically correct, but they can still mislead for two reasons. First, data may…

Databases · Computer Science 2019-09-13 Stavros Sintos , Pankaj K. Agarwal , Jun Yang