English
Related papers

Related papers: A Data-Driven Method for Automated Data Superposit…

200 papers

A ubiquitous feature of data of our era is their extra-large sizes and dimensions. Analyzing such high-dimensional data poses significant challenges, since the feature dimension is often much larger than the sample size. This thesis…

Statistics Theory · Mathematics 2025-09-11 Kai Yang

The paper considers the problem of establishing data support for the simplifying assumption (SA) in a bivariate conditional copula model. It is known that SA greatly simplifies the inference for a conditional copula model, but standard…

Methodology · Statistics 2019-09-30 Evgeny Levi , Radu V Craiu

The last decade witnessed an explosion in the availability of data for operations research applications. Motivated by this growing availability, we propose a novel schema for utilizing data to design uncertainty sets for robust optimization…

Optimization and Control · Mathematics 2014-11-25 Dimitris Bertsimas , Vishal Gupta , Nathan Kallus

We investigate a data-driven approach to constructing uncertainty sets for robust optimization problems, where the uncertain problem parameters are modeled as random variables whose joint probability distribution is not known. Relying only…

Optimization and Control · Mathematics 2020-09-22 Polina Alexeenko , Eilyan Bitar

Empirical process theory for i.i.d. observations has emerged as a ubiquitous tool for understanding the generalization properties of various statistical problems. However, in many applications where the data exhibit temporal dependencies…

Statistics Theory · Mathematics 2024-01-18 Nabarun Deb , Debarghya Mukherjee

Subsampling algorithms for various parametric regression models with massive data have been extensively investigated in recent years. However, all existing studies on subsampling heavily rely on clean massive data. In practical…

Statistics Theory · Mathematics 2025-06-11 Jiangshan Ju , Mingqiu Wang , Shengli Zhao

Statistical matching is an effective method for estimating causal effects in which treated units are paired with control units with ``similar'' values of confounding covariates prior to performing estimation. In this way, matching helps…

Methodology · Statistics 2023-09-13 Sanjeewani Weerasingha , Michael J. Higgins

Data assimilation (DA) aims at forecasting the state of a dynamical system by combining a mathematical representation of the system with noisy observations taking into account their uncertainties. State of the art methods are based on the…

Machine Learning · Computer Science 2023-05-26 Pierre Boudier , Anthony Fillion , Serge Gratton , Selime Gürol , Sixin Zhang

In a landscape where scientific discovery is increasingly driven by data, the integration of machine learning (ML) with traditional scientific methodologies has emerged as a transformative approach. This paper introduces a novel,…

Machine Learning · Computer Science 2024-06-26 Yunjin Tong

Circular and non-flat data distributions are prevalent across diverse domains of data science, yet their specific geometric structures often remain underutilized in machine learning frameworks. A principled approach to accounting for the…

Methodology · Statistics 2025-09-25 Thibault de Surrel , Fabien Lotte , Sylvain Chevallier , Florian Yger

Variable selection for high-dimensional, highly correlated data has long been a challenging problem, often yielding unstable and unreliable models. We propose a resample-aggregate framework that exploits diffusion models' ability to…

Methodology · Statistics 2025-08-20 Minjie Wang , Xiaotong Shen , Wei Pan

This paper considers the ill-posed data assimilation problem associated with hyperbolic/parabolic systems describing 2D coupled sound and heat flow. Given hypothetical data at time T > 0, that may not correspond to an actual solution of the…

Numerical Analysis · Mathematics 2025-01-28 Alfred S. Carasso

Traditional statistical and machine learning methods typically assume that the training and test data follow the same distribution. However, this assumption is frequently violated in real-world applications, where the training data in the…

Methodology · Statistics 2025-07-08 Hanxuan Ye , Hongzhe Li

A powerful tool for the analysis of nonrandomized observational studies has been the potential outcomes model. Utilization of this framework allows analysts to estimate average treatment effects. This article considers the situation in…

Statistics Theory · Mathematics 2019-05-31 Debashis Ghosh , Efrén Cruz-Cortés

We study the properties of conformal prediction for network data under various sampling mechanisms that commonly arise in practice but often result in a non-representative sample of nodes. We interpret these sampling mechanisms as selection…

Statistics Theory · Mathematics 2023-07-14 Robert Lunde

We present a convolution-based data assimilation method tailored to neuronal electrophysiology, addressing the limitations of traditional value-based synchronization approaches. While conventional methods rely on nudging terms and pointwise…

Neurons and Cognition · Quantitative Biology 2025-06-16 Dawei Li , Henry D. I. Abarbanel

An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable…

Machine learning is offering powerful new tools for the development and discovery of reduced models of nonlinear, multiscale plasma dynamics from the data of first-principles kinetic simulations. However, ensuring the physical consistency…

Plasma Physics · Physics 2026-02-25 Madox C. McGrae-Menge , Jacob R. Pierce , Frederico Fiuza , E. Paulo Alves

A key challenge in spatial statistics is the analysis for massive spatially-referenced data sets. Such analyses often proceed from Gaussian process specifications that can produce rich and robust inference, but involve dense covariance…

Methodology · Statistics 2019-07-25 Shinichiro Shirota , Andrew O. Finley , Bruce D. Cook , Sudipto Banerjee

Nonparametric regression for massive numbers of samples (n) and features (p) is an increasingly important problem. In big n settings, a common strategy is to partition the feature space, and then separately apply simple models to each…

Machine Learning · Statistics 2014-06-10 Rajarshi Guhaniyogi , David B. Dunson