Related papers: The Filtration of the split-words process
A. Vershik discovered that filtrations indexed by the non-positive integers may have a paradoxical asymptotic behaviour near the time $-\infty$, called non-standardness. For example, two dyadic filtrations with trivial tail $\sigma$-field…
The notion of a homogeneous standard filtration of $\sigma$-algebras was introduced by the author in 1970. The main theorem asserted that a homogeneous filtration is standard, i.e., generated by a sequence of independent random variables,…
We derive a practical standardness criterion for the filtration generated by a monotonic Markov process. This criterion is applied to show standardness of some adic filtrations.
We study the standard property of the natural filtration associated to a 0--1 valued stationary process. In our main result we show that if the process has summable memory decay, then the associated filtration is standard. We prove it by…
Let $X$ be a stationary process with finite state-space $A$. Bressaud et al. recently provided a sufficient condition for the natural filtration of $X$ to be standard when $A$ has size 2. Their condition involves the conditional laws…
The solution of the continuous time filtering problem can be represented as a ratio of two expectations of certain functionals of the signal process that are parametrized by the observation path. We introduce a class of discretization…
A filtration of a formal language L by a sequence s maps L to the set of words formed by taking the letters of words of L indexed only by s. We consider the languages resulting from filtering by all arithmetic progressions. If L is regular,…
Separation is a classical problem asking whether, given two sets belonging to some class, it is possible to separate them by a set from a smaller class. We discuss the separation problem for regular languages. We give a Ptime algorithm to…
Thanks to the nonstandard formalization of fast oscillating functions, due to P. Cartier and Y. Perrin, an appropriate mathematical framework is derived for new non-asymptotic estimation techniques, which do not necessitate any statistical…
Gorman and Bedrick (2019) argued for using random splits rather than standard splits in NLP experiments. We argue that random splits, like standard splits, lead to overly optimistic performance estimates. We can also split data in biased or…
The solution of the continuous time filtering problem can be represented as a ratio of two expectations of certain functionals of the signal process that are parametrized by the observation path. We introduce a new time discretisation of…
Given two languages, a separator is a third language that contains the first one and is disjoint from the second one. We investigate the following decision problem: given two regular input languages of finite words, decide whether there…
The generalized filtered method of moments was developed in the recent papers by Alomari et al., 2020, and Ayache et al., 2022. It used functional data obtained from continuously sampled cyclic long-memory stochastic processes to…
The Split and Rephrase (SPRP) task, which consists in splitting complex sentences into a sequence of shorter grammatical sentences, while preserving the original meaning, can facilitate the processing of complex texts for humans and…
We offer a natural and extensible measure-theoretic treatment of missingness at random. Within the standard missing data framework, we give a novel characterisation of the observed data as a stopping-set sigma algebra. We demonstrate that…
Often, when analyzing the behaviour of systems modelled as context-free languages, we wish to know if two languages overlap. To this end, we present an effective semi-decision procedure for regular separability of context-free languages,…
In this paper we consider the normalized lengths of the factors of some factorizations of random words. First, for the \emph{Lyndon factorization} of finite random words with $n$ independent letters drawn from a finite or infinite totally…
In this paper we study progressive filtration expansions with random times. We show how semimartingale decompositions in the expanded filtration can be obtained using a natural link between progressive and initial expansions. The link is,…
When expanding a filtration with a stochastic process it is easily possible for semimartingale no longer to remain semimartingales in the enlarged filtration. Y. Kchia and P. Protter indicated a way to avoid this pitfall in 2015, but they…
To mitigate the problem of having to traverse over the full vocabulary in the softmax normalization of a neural language model, sampling-based training criteria are proposed and investigated in the context of large vocabulary word-based…