English
Related papers

Related papers: A Two Parameters Equation for Word Rank-Frequency …

200 papers

Given a 2-SAT formula $F$ consisting of $n$ variables and $\cn$ random clauses, what is the largest number of clauses $\max F$ satisfiable by a single assignment of the variables? We bound the answer away from the trivial bounds of…

Combinatorics · Mathematics 2016-09-07 Don Coppersmith , David Gamarnik , Mohammad Hajiaghayi , Gregory B. Sorkin

Search techniques make use of elementary information such as term frequencies and document lengths in computation of similarity weighting. They can also exploit richer statistics, in particular the number of documents in which any two terms…

Information Retrieval · Computer Science 2020-07-20 Bodo Billerbeck , Justin Zobel , Nicholas Lester , Nick Craswell

An important body of quantitative linguistics is constituted by a series of statistical laws about language usage. Despite the importance of these linguistic laws, some of them are poorly formulated, and, more importantly, there is no…

Physics and Society · Physics 2020-11-09 Alvaro Corral , Isabel Serra

In the task of information retrieval the term relevance is taken to mean formal conformity of a document given by the retrieval system to user's information query. As a rule, the documents found by the retrieval system should be submitted…

Computation and Language · Computer Science 2007-10-02 S. Braichevsky , D. Lande , A. Snarskii

One of the main challenges in ranking is embedding the query and document pairs into a joint feature space, which can then be fed to a learning-to-rank algorithm. To achieve this representation, the conventional state of the art approaches…

Computation and Language · Computer Science 2018-08-09 Dana Sagi , Tzoof Avny , Kira Radinsky , Eugene Agichtein

This paper describes the probabilistic behaviour of a random Sturmian word. It performs the probabilistic analysis of the recurrence function which can be viewed as a waiting time to discover all the factors of length $n$ of the Sturmian…

Discrete Mathematics · Computer Science 2016-10-06 Pablo Rotondo , Brigitte Vallee

The monitoring of event frequencies can be used to recognize behavioral anomalies, to identify trends, and to deduce or discard hypotheses about the underlying system. For example, the performance of a web server may be monitored based on…

Logic in Computer Science · Computer Science 2020-01-13 Thomas Ferrère , Thomas A. Henzinger , Bernhard Kragl

We design alignment-free techniques for comparing a sequence or word, called a target, against a set of words, called a reference. A target-specific factor of a target $T$ against a reference $R$ is a factor $w$ of a word in $T$ which is…

Data Structures and Algorithms · Computer Science 2023-04-07 Marie-Pierre Béal , Maxime Crochemore

A simple reaction-diffusion-advection equation is proposed in a dichotomous tree network to discuss an optimal network. An optimal size ratio r is evaluated by the principle of maximization of total reaction rate. In the case of…

Statistical Mechanics · Physics 2015-06-23 Hidetsugu Sakaguchi

It is well known that a best rank-$R$ approximation of order-3 tensors may not exist for $R\ge 2$. A best rank-$(R,R,R)$ approximation always exists, however, and is also a best rank-$R$ approximation when it has rank (at most) $R$. For…

Algebraic Geometry · Mathematics 2016-09-23 Alwin Stegeman , Shmuel Friedland

We prove that any total boolean function of rank $r$ can be computed by a deterministic communication protocol of complexity $O(\sqrt{r} \cdot \log(r))$. Equivalently, any graph whose adjacency matrix has rank $r$ has chromatic number at…

Computational Complexity · Computer Science 2013-10-09 Shachar Lovett

The inverse relationship between the length of a word and the frequency of its use, first identified by G.K. Zipf in 1935, is a classic empirical law that holds across a wide range of human languages. We demonstrate that length is one…

Computation and Language · Computer Science 2017-06-02 Stephan C. Meylan , Thomas L. Griffiths

We study the distribution and the popularity of some patterns in $k$-ary faro words, i.e. words over the alphabet $\{1, 2, \ldots, k\}$ obtained by interlacing the letters of two nondecreasing words of lengths differing by at most one. We…

Combinatorics · Mathematics 2021-05-19 Jean-Luc Baril , Alexander Burstein , Sergey Kirgizov

A finite word $w$ with $\vert w\vert=n$ contains at most $n+1$ distinct palindromic factors. If the bound $n+1$ is attained, the word $w$ is called \emph{rich}. Let $\Factor(w)$ be the set of factors of the word $w$. It is known that there…

Combinatorics · Mathematics 2019-09-06 Josef Rukavicka

We describe a new method that is both physically explicable and quantitatively accurate in describing the multifractal characteristics of intermittent events based on groupings of rank-ordered fluctuations. The generic nature of such…

Astrophysics · Physics 2009-06-23 Tom Chang , Cheng-chin Wu

We develop fixed-point algorithms for the approximation of structured matrices with rank penalties. In particular we use these fixed-point algorithms for making approximations by sums of exponentials, or frequency estimation. For the basic…

Numerical Analysis · Mathematics 2016-01-07 Fredrik Andersson , Marcus Carlsson

This work is motivated by a hand-collected data set from one of the largest Internet portals in Korea. This data set records the top 30 most frequently discussed stocks on its on-line message board. The frequencies are considered to measure…

Methodology · Statistics 2017-01-05 Yuneung Kim , Johan Lim , Young-Geun Choi , Sujung Choi , Do Hwan Park

For words of length n, generated by independent geometric random variables, we consider the average value and the average position of the r-th left-to-right maximum, for fixed r and large n.

Combinatorics · Mathematics 2007-05-23 Arnold Knopfmacher , Helmut Prodinger

We study 4 problems in string matching, namely, regular expression matching, approximate regular expression matching, string edit distance, and subsequence indexing, on a standard word RAM model of computation that allows logarithmic-sized…

Data Structures and Algorithms · Computer Science 2008-09-22 Philip Bille , Martin Farach-Colton

We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…

Information Theory · Computer Science 2012-09-25 R. Lambiotte , M. Ausloos , M. Thelwall
‹ Prev 1 4 5 6 7 8 10 Next ›