中文
相关论文

相关论文: Statistical Integration of Heterogeneous Data with…

200 篇论文

The dramatic growth of big datasets presents a new challenge to data storage and analysis. Data reduction, or subsampling, that extracts useful information from datasets is a crucial step in big data analysis. We propose an orthogonal…

统计方法学 · 统计学 2021-06-01 Lin Wang , Jake Elmstedt , Weng Kee Wong , Hongquan Xu

Occupancy models involve both the probability a site is occupied and the probability occupancy is detected. The homogeneous occupancy model, where the occupancy and detection probabilities are the same at each site, admits an orthogonal…

统计方法学 · 统计学 2018-08-10 N. Karavarsamis , R. M. Huggins

In systems biology, it is becoming increasingly common to measure biochemical entities at different levels of the same biological system. Hence, data fusion problems are abundant in the life sciences. With the availability of a multitude of…

This paper proposes the asymmetric linear double autoregression, which jointly models the conditional mean and conditional heteroscedasticity characterized by asymmetric effects. A sufficient condition is established for the existence of a…

统计方法学 · 统计学 2021-04-22 Songhua Tan , Qianqian Zhu

Partial least squares regression---or PLS---is a multivariate method in which models are estimated using either the SIMPLS or NIPALS algorithm. PLS regression has been extensively used in applied research because of its effectiveness in…

统计方法学 · 统计学 2019-11-12 Titin Agustin Nengsih , Frédéric Bertrand , Myriam Maumy-Bertrand , Nicolas Meyer

We propose a novel method for multiple clustering that assumes a co-clustering structure (partitions in both rows and columns of the data matrix) in each view. The new method is applicable to high-dimensional data. It is based on a…

Assume that a finite set of points is randomly sampled from a subspace of a metric space. Recent advances in computational topology have provided several approaches to recovering the geometric and topological properties of the underlying…

代数拓扑 · 数学 2021-01-29 Peter Bubenik , Peter T. Kim

We present a generative learning framework for probabilistic sampling based on an extension of the Probabilistic Learning on Manifolds (PLoM) approach, which is designed to generate statistically consistent realizations of a random vector…

机器学习 · 统计学 2025-06-04 Dimitris G Giovanis , Nikolaos Evangelou , Ioannis G Kevrekidis , Roger G Ghanem

The rapid growth of high-dimensional datasets across various scientific domains has created a pressing need for new statistical methods to compare distributions supported on their underlying structures. Assessing similarity between datasets…

统计理论 · 数学 2025-11-27 Hongrui Chen , Rong Ma

There has been an intense recent activity in embedding of very high dimensional and nonlinear data structures, much of it in the data science and machine learning literature. We survey this activity in four parts. In the first part we cover…

机器学习 · 统计学 2022-09-01 Dag Tjøstheim , Martin Jullum , Anders Løland

In time series analysis, statistics based on collections of estimators computed from sub-samples play a crucial role in an increasing variety of important applications. Proving results about the joint asymptotic distribution of such…

统计理论 · 数学 2013-05-27 Stanislav Volgushev , Xiaofeng Shao

Partial Least Squares (PLS) regression emerged as an alternative to ordinary least squares for addressing multicollinearity in a wide range of scientific applications. As multidimensional tensor data is becoming more widespread, tensor…

统计方法学 · 统计学 2024-10-11 Kwangmoon Park , Sündüz Keleş

A significant hurdle for analyzing large sample data is the lack of effective statistical computing and inference methods. An emerging powerful approach for analyzing large sample data is subsampling, by which one takes a random subsample…

统计方法学 · 统计学 2015-11-24 Rong Zhu , Ping Ma , Michael W. Mahoney , Bin Yu

In biomedical research, to obtain more accurate prediction results from a target study, leveraging information from multiple similar source studies is proved to be useful. However, in many biomedical applications based on real-world data,…

统计方法学 · 统计学 2025-12-29 Xiaokang Liu , Jie Hu , Naimin Jing , Yang Ning , Cheng Yong Tang , Runze Li , Yong Chen

In this paper, we present an information-theoretic method for clustering mixed-type data, that is, data consisting of both continuous and categorical variables. The proposed approach extends the Information Bottleneck principle to…

统计方法学 · 统计学 2026-02-02 Efthymios Costa , Ioanna Papatsouma , Angelos Markos

Generalized linear models (GLMs) -- such as logistic regression, Poisson regression, and robust regression -- provide interpretable models for diverse data types. Probabilistic approaches, particularly Bayesian ones, allow coherent…

统计计算 · 统计学 2018-12-19 Jonathan H. Huggins , Ryan P. Adams , Tamara Broderick

Nonuniform subsampling methods are effective to reduce computational burden and maintain estimation efficiency for massive data. Existing methods mostly focus on subsampling with replacement due to its high computational efficiency. If the…

统计方法学 · 统计学 2021-07-06 Jun Yu , HaiYing Wang , Mingyao Ai , Huiming Zhang

A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be…

机器学习 · 计算机科学 2019-09-12 Jonas Mueller , Alex Smola

The PARAFAC2 is a multimodal factor analysis model suitable for analyzing multi-way data when one of the modes has incomparable observation units, for example because of differences in signal sampling or batch sizes. A fully probabilistic…

Integrating probability and non-probability samples is increasingly important, yet unknown sampling mechanisms in non-probability sources complicate identification and efficient estimation. We develop semiparametric theory for dual-frame…

统计方法学 · 统计学 2026-01-14 Kosuke Morikawa , Jae Kwang Kim