English
Related papers

Related papers: The correlation dimension of differenced data

200 papers

Time series data that are not measured at regular intervals are commonly discretized as a preprocessing step. For example, data about customer arrival times might be simplified by summing the number of arrivals within hourly intervals,…

Machine Learning · Statistics 2018-10-09 Peter Schulam , Suchi Saria

Inferring linear dependence between time series is central to our understanding of natural and artificial systems. Unfortunately, the hypothesis tests that are used to determine statistically significant directed or multivariate…

Methodology · Statistics 2021-02-24 Oliver M. Cliff , Leonardo Novelli , Ben D. Fulcher , James M. Shine , Joseph T. Lizier

In time-series analysis, the term "lead-lag effect" is used to describe a delayed effect on a given time series caused by another time series. lead-lag effects are ubiquitous in practice and are specifically critical in formulating…

Statistical Finance · Quantitative Finance 2020-02-04 Katsuya Ito , Kei Nakagawa

Time series are high-dimensional and complex data objects, making their efficient search and indexing a longstanding challenge in data mining. Building on a recently introduced similarity measure, namely Multiscale Dubuc Distance (MDD),…

Machine Learning · Computer Science 2025-10-28 Azim Ahmadzadeh , Mahsa Khazaei , Elaina Rohlfing

We consider the problem of surrogate sufficient dimension reduction, that is, estimating the central subspace of a regression model, when the covariates are contaminated by measurement error. When no measurement error is present, a…

Methodology · Statistics 2023-10-24 Linh H. Nghiem , Francis K. C. Hui , Samuel Mueller , A. H. Welsh

Regression Discontinuity Design (RDD) is a popular framework for estimating a causal effect in settings where treatment is assigned if an observed covariate exceeds a fixed threshold. We consider estimation and inference in the common…

Statistics Theory · Mathematics 2025-04-16 Kevin Tao , Y. Samuel Wang , David Ruppert

Surrogate models are used to reduce the burden of expensive-to-evaluate objective functions in optimization. By creating models which map genomes to objective values, these models can estimate the performance of unknown inputs, and so be…

Neural and Evolutionary Computing · Computer Science 2019-07-17 Alexander Hagg , Martin Zaefferer , Jörg Stork , Adam Gaier

Pseudorandom sequences are used extensively in communications and remote sensing. Correlation provides one measure of pseudorandomness, and low correlation is an important factor determining the performance of digital sequences in…

Information Theory · Computer Science 2018-06-14 Daniel J. Katz

In complex large-scale systems such as climate, important effects are caused by a combination of confounding processes that are not fully observable. The identification of sources from observations of system state is vital for attribution…

Machine Learning · Statistics 2023-03-22 Joseph Hart , Mamikon Gulian , Indu Manickam , Laura Swiler

A common challenge in computer experiments and related fields is to efficiently explore the input space using a small number of samples, i.e., the experimental design problem. Much of the recent focus in the computer experiment literature,…

Methodology · Statistics 2019-07-01 Boya Zhang , D. Austin Cole , Robert B. Gramacy

Recently, the quantification of errors in the stochastic homogenization of divergence-form operators has witnessed important progress. Our aim now is to go beyond error bounds, and give precise descriptions of the effect of the randomness,…

Analysis of PDEs · Mathematics 2016-09-29 Jean-Christophe Mourrat , Felix Otto

We consider the problem of approximating sums of high-dimensional stationary time series by Gaussian vectors, using the framework of functional dependence measure. The validity of the Gaussian approximation depends on the sample size $n$,…

Statistics Theory · Mathematics 2015-08-31 Danna Zhang , Wei Biao Wu

Cause-effect relationships are typically evaluated by comparing outcome responses to binary treatment values, representing two arms of a hypothetical randomized controlled trial. However, in certain applications, treatments of interest are…

Methodology · Statistics 2022-06-15 Razieh Nabi , Todd McNutt , Ilya Shpitser

Multifidelity surrogate modelling combines data of varying accuracy and cost from different sources. It strategically uses low-fidelity models for rapid evaluations, saving computational resources, and high-fidelity models for detailed…

Machine Learning · Computer Science 2024-04-24 Daniel N Wilke

Surrogate markers are often used in clinical trials to evaluate treatment effects when primary outcomes are costly, invasive, or take a long time to observe. However, reliance on surrogates can lead to the surrogate paradox, where a…

Methodology · Statistics 2025-06-17 Emily Hsiao , Lu Tian , Layla Parast

Accurate estimation for extent of cross{sectional dependence in large panel data analysis is paramount to further statistical analysis on the data under study. Grouping more data with weak relations (cross{sectional dependence) together…

Econometrics · Economics 2019-04-16 Jiti Gao , Guangming Pan , Yanrong Yang , Bo Zhang

This paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test…

Computation and Language · Computer Science 2016-09-28 Magnus Sahlgren , Alessandro Lenci

In the presence of weak overall correlation, it may be useful to investigate if the correlation is significantly and substantially more pronounced over a subpopulation. Two different testing procedures are compared. Both are based on the…

Machine Learning · Statistics 2015-04-22 Stephen Bamattre , Rex Hu , Joseph S. Verducci

Many real-world multivariate time series are collected from a network of physical objects embedded with software, electronics, and sensors. The quasi-periodic signals generated by these objects often follow a similar repetitive and periodic…

Machine Learning · Computer Science 2025-06-23 Kai Yang , Shaoyu Dou , Pan Luo , Xin Wang , H. Vincent Poor

The Consistency property between surrogate losses and evaluation metrics has been extensively studied to ensure that minimizing a loss leads to metric optimality. However, the direct relationship between different evaluation metrics remains…

Machine Learning · Computer Science 2026-03-10 Yuanhao Pu , Defu Lian , Enhong Chen