English
Related papers

Related papers: A Two Parameters Equation for Word Rank-Frequency …

200 papers

Assume in a sample of size M one finds M_i representatives of species i with i=1...N^*. The normalized frequency p^*_i=M_i/M, based on the finite sample, may deviate considerably from the true probabilities p_i. We propose a method to infer…

Other Quantitative Biology · Quantitative Biology 2007-05-23 Thorsten Poeschel , Werner Ebeling , Cornelius Froemmel , Rosa Ramirez

This paper considers the use of recently proposed optimal transport-based multivariate test statistics, namely rank energy and its variant the soft rank energy derived from entropically regularized optimal transport, for the unsupervised…

Machine Learning · Statistics 2023-02-17 Matthew Werenski , Shoaib Bin Masud , James M. Murphy , Shuchin Aeron

Recommendations based on behavioral data may be faced with ambiguous statistical evidence. We consider the case of association rules, relevant e.g.~for query and product recommendations. For example: Suppose that a customer belongs to…

Databases · Computer Science 2015-01-12 Rasmus Pagh , Morten Stöckel

Prior to recent successes using neural networks, term frequency-inverse document frequency (tf-idf) was clearly regarded as the best choice for identifying documents related to a query. We provide a different score, aver, and observe, on a…

Information Retrieval · Computer Science 2025-11-10 Anthony Gamst , Lawrence Wilson

Regular expressions with backreferences (regex, for short), as supported by most modern libraries for regular expression matching, have an NP-complete matching problem. We define a complexity parameter of regex, called active variable…

Formal Languages and Automata Theory · Computer Science 2024-02-09 Markus L. Schmid

Let $x$ and $y$ be words. We consider the languages whose words $z$ are those for which the numbers of occurrences of $x$ and $y$, as subwords of $z$, are the same (resp., the number of $x$'s is less than the number of $y$'s, resp., is less…

Formal Languages and Automata Theory · Computer Science 2018-06-22 Charles J. Colbourn , Ryan E. Dougherty , Thomas F. Lidbetter , Jeffrey Shallit

Let $\mathscr{L}$ be a recursive language. Let $S(\mathscr{L})$ be the set of $\mathscr{L}$-structures with domain $\omega$. Let $\Phi : {}^\omega 2 \rightarrow S(\mathscr{L})$ be a $\Delta_1^1$ function with the property that for all $x,y…

Logic · Mathematics 2017-12-05 William Chan , Matthew Harrison-Trainor , Andrew Marks

The rank of a graph is defined to be the rank of its adjacency matrix. A graph is called reduced if it has no isolated vertices and no two vertices with the same set of neighbors. Akbari, Cameron, and Khosrovshahi conjectured that the…

Combinatorics · Mathematics 2014-04-29 E. Ghorbani , A. Mohammadian , B. Tayfeh-Rezaie

Given a finite word $w$ over a finite alphabet $V$, consider the graph with vertex set $V$ and with an edge between two elements of $V$ if and only if the two elements alternate in the word $w$. Such a graph is said to be word-representable…

Combinatorics · Mathematics 2021-01-15 Marisa Gaetz , Caleb Ji

Selecting the best items in a dataset is a common task in data exploration. However, the concept of "best" lies in the eyes of the beholder: different users may consider different attributes more important, and hence arrive at different…

Databases · Computer Science 2023-04-27 Abolfazl Asudeh , Azade Nazi , Nan Zhang , Gautam Das , H. V. Jagadish

Ranking is a key aspect of many applications, such as information retrieval, question answering, ad placement and recommender systems. Learning to rank has the goal of estimating a ranking model automatically from training data. In…

Information Retrieval · Computer Science 2015-02-10 Truyen Tran , Dinh Phung , Svetha Venkatesh

It is well-known that checking whether a given string $w$ matches a given regular expression $r$ can be done in quadratic time $O(|w|\cdot |r|)$ and that this cannot be improved to a truly subquadratic running time of $O((|w|\cdot…

Data Structures and Algorithms · Computer Science 2025-08-21 Antoine Amarilli , Florin Manea , Tina Ringleb , Markus L. Schmid

In this paper we initiate the study of computing a maximal (not necessarily maximum) repeating pattern in a single input string, where the corresponding problems have been studied (e.g., a maximal common subsequence) only in two or more…

Data Structures and Algorithms · Computer Science 2026-01-21 Mingyang Gong , Adiesha Liyanage , Braeden Sopp , Binhai Zhu

We prove that, for any pure morphic word $w$, if the frequencies of all letters in $w$ exist, then the frequencies of all factors in $w$ exist as well. This result answers a question of Saari in his doctoral thesis.

Combinatorics · Mathematics 2024-05-30 Shuo Li

We consider random walks on the set of all words over a finite alphabet such that in each step only the last two letters of the current word may be modified and only one letter may be adjoined or deleted. We assume that the transition…

Probability · Mathematics 2008-07-16 Lorenz A. Gilch

This paper proposes an algorithm to improve the calculation of confidence measure for spoken term detection (STD). Given an input query term, the algorithm first calculates a measurement named document ranking weight for each document in…

Computation and Language · Computer Science 2015-09-11 Quan Liu , Wu Guo , Zhen-Hua Ling

A finite word $w$ is called \emph{rich} if it contains $\vert w\vert+1$ distinct palindromic factors including the empty word. Let $q\geq 2$ be the size of the alphabet. Let $R(n)$ be the number of rich words of length $n$. Let $d>1$ be a…

Combinatorics · Mathematics 2022-12-20 Josef Rukavicka

One of the major sources of trending news, events and opinion in the current age is micro blogging. Twitter, being one of them, is extensively used to mine data about public responses and event updates. This paper intends to propose methods…

Social and Information Networks · Computer Science 2015-06-22 Rishabh Jain , Abhishek B. S. , Satvik Jagannath

Zipf's law on word frequency is observed in English, French, Spanish, Italian, and so on, yet it does not hold for Chinese, Japanese or Korean characters. A model for writing process is proposed to explain the above difference, which takes…

Data Analysis, Statistics and Probability · Physics 2013-05-03 Linyuan Lu , Zi-Ke Zhang , Tao Zhou

The frequency of the preferred order for a noun phrase formed by demonstrative, numeral, adjective and noun has received significant attention over the last two decades. We investigate the actual distribution of the 24 possible orders.…

Computation and Language · Computer Science 2026-01-23 Ramon Ferrer-i-Cancho