English
Related papers

Related papers: Customized determination of stop words using Rando…

200 papers

Any positive word comprised of random sequence of tokens form a finite alphabet can be reduced (without change of length) using an appropriate size Braid group relationships. Surprisingly the Braid relations dramatically reduce the…

Computational Complexity · Computer Science 2013-08-20 Dara O Shayda

In this paper, we consider robust control using randomized algorithms. We extend the existing order statistics distribution theory to the general case in which the distribution of population is not assumed to be continuous and the order…

Optimization and Control · Mathematics 2008-05-13 Xinjia Chen , Kemin Zhou

Let $T\$ be a stopping time associated with a sequence of independent random variables $Z_{1},Z_{2},...$ . By applying a suitable change in the probability measure we present relations between the moment or probability generating functions…

Statistics Theory · Mathematics 2011-06-28 M. V. Boutsikas , A. C. Rakitzis , D. L. Antzoulakos

Learning word embeddings using distributional information is a task that has been studied by many researchers, and a lot of studies are reported in the literature. On the contrary, less studies were done for the case of multiple languages.…

Computation and Language · Computer Science 2020-04-15 Marco Berlot , Evan Kaplan

Deterministic automata have been traditionally studied through the point of view of language equivalence, but another perspective is given by the canonical notion of shortest-distinguishing-word distance quantifying the of states.…

Logic in Computer Science · Computer Science 2024-04-23 Wojciech Różowski

Existing training criteria in automatic speech recognition(ASR) permit the model to freely explore more than one time alignments between the feature and label sequences. In this paper, we use entropy to measure a model's uncertainty, i.e.…

Computation and Language · Computer Science 2022-12-26 Ehsan Variani , Ke Wu , David Rybach , Cyril Allauzen , Michael Riley

The analysis of strings of $n$ random variables with geometric distribution has recently attracted renewed interest: Archibald et al. consider the number of distinct adjacent pairs in geometrically distributed words. They obtain the…

Probability · Mathematics 2024-02-14 Guy Louchard , Werner Schachinger , Mark Daniel Ward

In this work we seek clusters of genomic words in human DNA by studying their inter-word lag distributions. Due to the particularly spiked nature of these histograms, a clustering procedure is proposed that first decomposes each…

Applications · Statistics 2021-01-13 Ana Helena Tavares , Jakob Raymaekers , Peter J. Rousseeuw , Paula Brito , Vera Afreixo

Detection of the number of signals corrupted by high-dimensional noise is a fundamental problem in signal processing and statistics. This paper focuses on a general setting where the high-dimensional noise has an unknown complicated…

Statistics Theory · Mathematics 2022-05-16 Xiucai Ding , Fan Yang

We present a new adaptive algorithm for learning discrete distributions under distribution drift. In this setting, we observe a sequence of independent samples from a discrete distribution that is changing over time, and the goal is to…

Machine Learning · Computer Science 2024-03-11 Alessio Mazzetto

Sampling-based motion planning algorithms are widely used in robotics because they are very effective in high-dimensional spaces. However, the success rate and quality of the solutions are determined by an adequate selection of their…

We investigate invariants for random elements of different hyperbolic groups. We provide a method, using Cayley graphs of groups, to compute the probability distribution of the minimal length of a random word, and explicitly compute the…

Mathematical Physics · Physics 2007-05-23 Sergei Nechaev , Raphael Voituriez

With generalizing the Brody distribution to include the Poisson, GOE and GUE limits and with employing the maximum likelihood estimation technique, the spectral statistics of different sequences were considered in the nearest neighbor…

Nuclear Theory · Physics 2012-10-18 M. A. Jafarizadeh , N. Fouladi , H. Sabri , B. R. Maleki

In this paper, we consider a network of processors aiming at cooperatively solving mixed-integer convex programs subject to uncertainty. Each node only knows a common cost function and its local uncertain constraint set. We propose a…

Optimization and Control · Mathematics 2022-07-19 Mohammadreza Chamanbaz , Giuseppe Notarstefano , Francesco Sasso , Roland Bouffanais

We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…

Information Theory · Computer Science 2012-09-25 R. Lambiotte , M. Ausloos , M. Thelwall

We describe a method for automatic word sense disambiguation using a text corpus and a machine-readable dictionary (MRD). The method is based on word similarity and context similarity measures. Words are considered similar if they appear in…

cmp-lg · Computer Science 2008-02-03 Yael Karov , Shimon Edelman

In this work, we unveil and study idiosyncrasies in Large Language Models (LLMs) -- unique patterns in their outputs that can be used to distinguish the models. To do so, we consider a simple classification task: given a particular text…

Computation and Language · Computer Science 2025-06-17 Mingjie Sun , Yida Yin , Zhiqiu Xu , J. Zico Kolter , Zhuang Liu

Cross-lingual or cross-domain correspondences play key roles in tasks ranging from machine translation to transfer learning. Recently, purely unsupervised methods operating on monolingual embeddings have become effective alignment tools.…

Computation and Language · Computer Science 2018-09-05 David Alvarez-Melis , Tommi S. Jaakkola

Numerous analyses of reading time (RT) data have been implemented -- all in an effort to better understand the cognitive processes driving reading comprehension. However, data measured on words at the end of a sentence -- or even at the end…

Computation and Language · Computer Science 2024-01-08 Clara Meister , Tiago Pimentel , Thomas Hikaru Clark , Ryan Cotterell , Roger Levy

We define and completely solve a content-based directed network whose nodes consist of random words and an adjacency rule involving perfect or approximate matches, for an alphabet with an arbitrary number of letters. The analytic expression…

Molecular Networks · Quantitative Biology 2007-05-23 Muhittin Mungan , Alkan Kabakcioglu , Duygu Balcan , Ayse Erzan