中文
相关论文

相关论文: A Common-Factor Approach for Multivariate Data Cle…

200 篇论文

Multimodal data, where different types of data are collected from the same subjects, are fast emerging in a large variety of scientific applications. Factor analysis is commonly used in integrative analysis of multimodal data, and is…

统计理论 · 数学 2021-03-31 Quefeng Li , Lexin Li

Data cleaning is often framed as a technical preprocessing step, yet in practice it relies heavily on human judgment. We report results from a controlled survey study in which participants performed error detection, data repair and…

数据库 · 计算机科学 2026-03-26 Hazim AbdElazim , Shadman Islam , Mostafa Milani

This paper proposes an imputation procedure that uses the factors estimated from a tall block along with the re-rotated loadings estimated from a wide block to impute missing values in a panel of data. Assuming that a strong factor…

计量经济学 · 经济学 2021-08-13 Jushan Bai , Serena Ng

The need for multimodal data integration arises naturally when multiple complementary sets of features are measured on the same sample. Under a dependent multifactor model, we develop a fully data-driven orchestrated approximate message…

统计方法学 · 统计学 2026-01-22 Sagnik Nandy , Zongming Ma

By creating networks of biochemical pathways, communities of micro-organisms are able to modulate the properties of their environment and even the metabolic processes within their hosts. Next-generation high-throughput sequencing has led to…

应用统计 · 统计学 2023-03-28 Molly G. Hayes , Morgan G. I. Langille , Hong Gu

Complex, high-dimensional data is used in a wide range of domains to explore problems and make decisions. Analysis of high-dimensional data, however, is vulnerable to the hidden influence of confounding variables, especially as users apply…

人机交互 · 计算机科学 2022-07-01 Smiti Kaul , David Borland , Nan Cao , David Gotz

Generalist imitation learning policies trained on large datasets show great promise for solving diverse manipulation tasks. However, to ensure generalization to different conditions, policies need to be trained with data collected across a…

Factor analysis is a flexible technique for assessment of multivariate dependence and codependence. Besides being an exploratory tool used to reduce the dimensionality of multivariate data, it allows estimation of common factors that often…

应用统计 · 统计学 2020-05-08 Vitor G. C. da Silva , Kelly C. M. Gonçalves , João B. M. Pereira

This paper presents a new statistical method for clustering step data, a popular form of health record data easily obtained from wearable devices. Since step data are high-dimensional and zero-inflated, classical methods such as K-means and…

统计方法学 · 统计学 2020-10-16 Wookyeong Song , Hee-Seok Oh , Yaeji Lim , Ying Kuen Cheung

Factor modeling is an essential tool for exploring intrinsic dependence structures among high-dimensional random variables. Much progress has been made for estimating the covariance matrix from a high-dimensional factor model. However, the…

统计理论 · 数学 2016-10-26 Quefeng Li , Guang Cheng , Jianqing Fan , Yuyan Wang

Multivariate spatially-oriented data sets are prevalent in the environmental and physical sciences. Scientists seek to jointly model multiple variables, each indexed by a spatial location, to capture any underlying spatial association for…

统计方法学 · 统计学 2021-08-19 Lu Zhang , Sudipto Banerjee

We propose modeling raw functional data as a mixture of a smooth function and a highdimensional factor component. The conventional approach to retrieving the smooth function from the raw data is through various smoothing techniques.…

统计方法学 · 统计学 2021-02-05 Yuan Gao , Han Lin Shang , Yanrong Yang

With the rapid development of the internet technology, dirty data are commonly observed in various real scenarios, e.g., owing to unreliable sensor reading, transmission and collection from heterogeneous sources. To deal with their negative…

数据库 · 计算机科学 2020-11-24 Yu Sun , Jian Zhang

The collection, transfer and integration of research information into different research Information systems can result in different data errors that can have a variety of negative effects on data quality. In order to detect errors at an…

数据库 · 计算机科学 2019-01-21 Otmane Azeroual , Gunter Saake , Mohammad Abuosba

Missing data are ubiquitous in the era of big data and, if inadequately handled, are known to lead to biased findings and have deleterious impact on data-driven decision makings. To mitigate its impact, many missing value imputation methods…

机器学习 · 计算机科学 2021-10-26 Yiliang Zhang , Qi Long

Data fusion, the process of combining observational and experimental data, can enable the identification of causal effects that would otherwise remain non-identifiable. Although identification algorithms have been developed for specific…

机器学习 · 统计学 2025-12-22 Otto Tabell , Santtu Tikka , Juha Karvanen

In the era of big data, the explosive growth of multi-source heterogeneous data offers many exciting challenges and opportunities for improving the inference of conditional average treatment effects. In this paper, we investigate…

机器学习 · 统计学 2022-11-02 Xinyu Li , Yilin Li , Qing Cui , Longfei Li , Jun Zhou

We propose a new method to impute missing values in mixed datasets. It is based on a principal components method, the factorial analysis for mixed data, which balances the influence of all the variables that are continuous and categorical…

应用统计 · 统计学 2013-02-20 Vincent Audigier , François Husson , Julie Josse

There is an implicit assumption in software testing that more diverse and varied test data is needed for effective testing and to achieve different types and levels of coverage. Generic approaches based on information theory to measure and…

软件工程 · 计算机科学 2017-09-19 Robert Feldt , Simon Poulding