English
Related papers

Related papers: Properties of maximum Lempel-Ziv complexity string…

200 papers

Frequent pattern mining is widely used to find ``important'' or ``interesting'' patterns in data. While it is not easy to mathematically define such patterns, maximal frequent patterns are promising candidates, as frequency is a natural…

Data Structures and Algorithms · Computer Science 2025-04-08 Giovanni Buzzega , Alessio Conte , Yasuaki Kobayashi , Kazuhiro Kurita , Giulia Punzi

For both the Lempel Ziv 77- and 78-factorization we propose algorithms generating the respective factorization using $(1+\epsilon) n \lg n + O(n)$ bits (for any positive constant $\epsilon \le 1$) working space (including the space for the…

Data Structures and Algorithms · Computer Science 2015-04-13 Johannes Fischer , Tomohiro I , Dominik Köppl

We introduce the first self-index based on the Lempel-Ziv 1977 compression format (LZ77). It is particularly competitive for highly repetitive text collections such as sequence databases of genomes of related species, software repositories,…

Data Structures and Algorithms · Computer Science 2011-01-24 Sebastian Kreft , Gonzalo Navarro

Maximal green sequences are important objects in representation theory, cluster algebras, and string theory. The two fundamental questions about maximal green sequences are whether a given algebra admits such sequences and, if so, does it…

Representation Theory · Mathematics 2020-10-30 Alexander Garver , Khrystyna Serhiyenko

In this paper, we study stability of $M$-compactness for $l^p$ sum of Banach spaces for $1\leq p<\infty$. We also obtain a characterization of $M$-compact sets in terms of statistically maximizing sequence, a notion which is weaker than a…

Functional Analysis · Mathematics 2020-08-20 Susmita Seal , Sumit Som , Sudeshna Basu , Lakshmi Kanta Dey

New classes of distance-constrained structures are introduced, namely string-node nets and meshes, a mesh being a string-node net for which the nodes are dense in the strings. Various construction schemes are given including the minimal…

Metric Geometry · Mathematics 2016-09-12 S. C. Power , B. Schulze

We observe a different type of complex solutions in the isotropic spin-1/2 Heisenberg chain starting from N=12, where the central rapidity of some of the odd-length strings becomes complex making not all the strings self-conjugate…

High Energy Physics - Theory · Physics 2015-02-12 Tetsuo Deguchi , Pulak Ranjan Giri

We report on a detailed numerical study of the evolution of semilocal string networks, based on the largest and most accurate field theory simulations of these objects to date. We focus on the large-scale network properties, confirming…

High Energy Physics - Phenomenology · Physics 2014-03-24 A. Achúcarro , A. Avgoustidis , A. M. M. Leite , A. Lopez-Eiguren , C. J. A. P. Martins , A. S. Nunes , J. Urrestilla

In this paper, we study for the first time the Diverse Longest Common Subsequences (LCSs) problem under Hamming distance. Given a set of a constant number of input strings, the problem asks to decide if there exists some subset $\mathcal X$…

Data Structures and Algorithms · Computer Science 2024-06-12 Yuto Shida , Giulia Punzi , Yasuaki Kobayashi , Takeaki Uno , Hiroki Arimura

String complexity is defined as the cardinality of a set of all distinct words (factors) of a given string. For two strings, we introduce the joint string complexity as the cardinality of a set of words that are common to both strings.…

Information Theory · Computer Science 2018-05-24 Philippe Jacquet , Dimitris Milioris , Wojciech Szpankowski

Adaptations of features commonly applied in the field of visual computing, co-occurrence matrix (COM) and run-length matrix (RLM), are proposed for the similarity computation of strings in general (words, phrases, codes and texts). The…

Machine Learning · Computer Science 2026-05-15 E. O. Rodrigues , D. Casanova , M. Teixeira , V. Pegorini , F. Favarim , E. Clua , A. Conci , Panos Liatsis

We study three fundamental statistical-learning problems: distribution estimation, property estimation, and property testing. We establish the profile maximum likelihood (PML) estimator as the first unified sample-optimal approach to a wide…

Machine Learning · Statistics 2019-07-12 Yi Hao , Alon Orlitsky

In this paper, we study the weight spectrum of linear codes with \emph{super-linear} field size and use the probabilistic method to show that for nearly all such codes, the corresponding weight spectrum is very close to that of a maximum…

Information Theory · Computer Science 2021-08-18 Ghurumuruhan Ganesan

Autoregressive language models (LMs) map token sequences to probabilities. The usual practice for computing the probability of any character string (e.g. English sentences) is to first transform it into a sequence of tokens that is scored…

Computation and Language · Computer Science 2023-07-03 Nadezhda Chirkova , Germán Kruszewski , Jos Rozen , Marc Dymetman

We discuss the problem of estimating the characteristic length scale $\xi_{\rm s}$, and hence the initial density, of a system of cosmic strings formed at a continuous, second-order phase transition in the early universe. In particular, we…

High Energy Physics - Phenomenology · Physics 2016-09-01 T. W. B. Kibble , Alexander Vilenkin

In most practical problems of classifier learning, the training data suffers from the label noise. Hence, it is important to understand how robust is a learning algorithm to such label noise. This paper presents some theoretical analysis to…

Machine Learning · Computer Science 2016-08-29 Aritra Ghosh , Naresh Manwani , P. S. Sastry

What can large language models learn? By definition, language models (LM) are distributions over strings. Therefore, an intuitive way of addressing the above question is to formalize it as a matter of learnability of classes of…

Computation and Language · Computer Science 2025-01-14 Nadav Borenstein , Anej Svete , Robin Chan , Josef Valvoda , Franz Nowak , Isabelle Augenstein , Eleanor Chodroff , Ryan Cotterell

This article gives a self-contained analysis of the performance of the Lempel-Ziv compression algorithm on (hidden) Markovian sources. Specifically we include a full proof of the assertion that the compression rate approaches the entropy…

Information Theory · Computer Science 2019-10-03 Madhu Sudan , David Xiang

Given two {0,1}-sequences X and Y of lengths m and n, respectively, we write L(X,Y) to denote the length of the longest common subsequence (LCS) of X and Y, and write L(m,n) to denote the expected value of L(X,Y) when X and Y are random…

Group Theory · Mathematics 2013-07-11 John D. Dixon

Detecting and measuring repetitiveness of strings is a problem that has been extensively studied in data compression and text indexing. However, when the data are structured in a non-linear way, like in the context of two-dimensional…

Data Structures and Algorithms · Computer Science 2024-04-11 Giuseppe Romana , Marinella Sciortino , Cristian Urbina