Related papers: Prediction of long memory processes on same-realis…
Deep Gaussian Processes learn probabilistic data representations for supervised learning by cascading multiple Gaussian Processes. While this model family promises flexible predictive distributions, exact inference is not tractable.…
The field of machine have seen rising applications of equivariance criterion. However, there is no systematic way to justify its usage, including why it works, whether there is an optimal solution and if so, what form it carries. In this…
This paper deals with the consistency of the least squares estimator of a convex regression function when the predictor is multidimensional. We characterize and discuss the computation of such an estimator via the solution of certain…
Successive Halving is a popular algorithm for hyperparameter optimization which allocates exponentially more resources to promising candidates. However, the algorithm typically relies on intermediate performance values to make resource…
Finitarily Markovian processes are those processes $\{X_n\}_{n=-\infty}^{\infty}$ for which there is a finite $K$ ($K = K(\{X_n\}_{n=-\infty}^0$) such that the conditional distribution of $X_1$ given the entire past is equal to the…
Iteration of randomly chosen quadratic maps defines a Markov process: X_{n+1}=\epsilon_{n+1}X_n(1-X_n), where \epsilon_n are i.i.d. with values in the parameter space [0,4] of quadratic maps F_{\theta}(x)=\theta x(1-x). Its study is of…
Empirical risk minimization is a standard principle for choosing algorithms in learning theory. In this paper we study the properties of empirical risk minimization for time series. The analysis is carried out in a general framework that…
In this work, we consider the deterministic optimization using random projections as a statistical estimation problem, where the squared distance between the predictions from the estimator and the true solution is the error metric. In…
A continuous-time regression model with a jointly strictly sub-Gaussian random noise is considered in the paper. Upper exponential bounds for probabilities of large deviations of the least squares estimator for the regression parameter are…
We consider the multilinear polynomial-form process \[X(n)=\sum_{1\le i_1<\ldots<i_k<\infty}a_{i_1}\ldots a_{i_k}\epsilon_{n-i_1}\ldots\epsilon_{n-i_k},\] obtained by applying a multilinear polynomial-form filter to i.i.d.\ sequence…
A local linear kernel estimator of the regression function x\mapsto g(x):=E[Y_i|X_i=x], x\in R^d, of a stationary (d+1)-dimensional spatial process {(Y_i,X_i),i\in Z^N} observed over a rectangular domain of the form I_n:={i=(i_1,...,i_N)\in…
Employing recent results of Robinson (2005) we consider the asymptotic properties of conditional-sum-of-squares (CSS) estimates of parametric models for stationary time series with long memory. CSS estimation has been considered as a rival…
Given a stationary first-order autoregressive process X_t (with lag-one correlation rho satisfying |rho|<1), we examine the Central Limit Theorem for (1/n)*ln |X_1...X_n| and compute variances to high precision. Given a nonstationary…
In this paper we propose the first non-parametric Bayesian model using Gaussian Processes to make inference on Poisson Point Processes without resorting to gridding the domain or to introducing latent thinning points. Unlike competing…
Previous analysis on forecasting theory either assume knowing the true parameters or assume the stationarity of the series. Not much are known on the forecasting theory for nonstationary process with estimated parameters. This paper…
We develop a Bayesian approach to learning from sequential data by using Gaussian processes (GPs) with so-called signature kernels as covariance functions. This allows to make sequences of different length comparable and to rely on strong…
We present a purely deep neural network-based approach for estimating long memory parameters of time series models that incorporate the phenomenon of long-range dependence. Parameters, such as the Hurst exponent, are critical in…
In this paper, we study finite-sample properties of the least squares estimator in first order autoregressive processes. By leveraging a result from decoupling theory, we derive upper bounds on the probability that the estimate deviates by…
We study the problem of list-decodable Gaussian mean estimation and the related problem of learning mixtures of separated spherical Gaussians. We develop a set of techniques that yield new efficient algorithms with significantly improved…
Causal Transformers are trained to predict the next token for a given context. While it is widely accepted that self-attention is crucial for encoding the causal structure of sequences, the precise underlying mechanism behind this…