Related papers: A Two Parameters Equation for Word Rank-Frequency …
Text classification is one of the most frequent tasks for processing textual data, facilitating among others research from large-scale datasets. Embeddings of different kinds have recently become the de facto standard as features used for…
This article characterizes the associativity of two-place functions $T: [0,1]^2\rightarrow [0,1]$ defined by $T(x,y)=f^{(-1)}(F(f(x),f(y)))$ where $F:[0,1]^2\rightarrow[0,1]$ is a triangular norm (even a triangular subnorm), $f:…
We study the rank of the instantaneous or spot covariance matrix $\Sigma_X(t)$ of a multidimensional continuous semi-martingale $X(t)$. Given high-frequency observations $X(i/n)$, $i=0,\ldots,n$, we test the null hypothesis…
We give a language-parametric solution to the problem of total correctness, by automatically reducing it to the problem of partial correctness, under the assumption that an expression whose value decreases with each program step in a…
Compositionality in language refers to how much the meaning of some phrase can be decomposed into the meaning of its constituents and the way these constituents are combined. Based on the premise that substitution by synonyms is…
Ranks estimated from data are uncertain and this poses a challenge in many applications. However, estimated ranks are deterministic functions of estimated parameters, so the uncertainty in the ranks must be determined by the uncertainty in…
In this paper, we propose a novel word-alignment-based method to solve the FAQ-based question answering task. First, we employ a neural network model to calculate question similarity, where the word alignment between two questions is used…
For $0<\delta <1$ a $\delta$-subrepetition in a word is a factor which exponent is less than~2 but is not less than $1+\delta$ (the exponent of the factor is the ratio of the factor length to its minimal period). The $\delta$-subrepetition…
Let $f\in \mathbb{R}[x_1,\ldots, x_k]$, for $k\ge 2$. For any finite sets $A_1,\ldots, A_k\subset \mathbb{R}$, consider the set $$ f(A_1,\ldots, A_k):=\{f(a_1,\ldots, a_k)\mid (a_1,\cdots,a_k)\in A_1\times\cdots \times A_k\}, $$ that is,…
We develop a static complexity analysis for a higher-order functional language with structural list recursion. The complexity of an expression is a pair consisting of a cost and a potential. The former is defined to be the size of the…
In the first part of the present work we consider periodically or quasiperiodically forced systems of the form $(d/dt)x = \epsilon f(x,t \omega )$, where $\epsilon\ll 1$, $\omega\in\mathbb{R}^d$ is a nonresonant vector of frequencies and…
We investigate conditions in order to decide whether a given sequence of real numbers represents expected maxima or expected ranges. The main result provides a novel necessary and sufficient condition, relating an expected maxima sequence…
According to the Probability Ranking Principle (PRP), ranking documents in decreasing order of their probability of relevance leads to an optimal document ranking for ad-hoc retrieval. The PRP holds when two conditions are met: [C1] the…
We investigate the following problem: given a sample of classified strings, find a first-order sentence of minimal quantifier rank that is consistent with the sample. We represent strings as successor string structures, that is, finite…
When aggregating preferences of agents via voting, two desirable goals are to incentivize agents to participate in the voting process and then identify outcomes that are Pareto efficient. We consider participation as formalized by Brandl,…
For $d\ge 1$, a word $w\in \{ 0,1\}^{\Z^d}$ is called balanced if there exists $M > 0$ such that for any two rectangles $R, R^{'}\subset\Z^d$ that are translates of each other, the number of occurrences of the symbol $1$ in $R$ and $R^{'}$…
We consider words as a network of interacting letters, and approximate the probability distribution of states taken on by this network. Despite the intuition that the rules of English spelling are highly combinatorial (and arbitrary), we…
The article attempts to find an algebraic formula describing the correlation coefficients between random variables and the principal components representing them. As a result of the analysis, starting from selected statistics relating to…
For bounded datasets such as the TREC Web Track (WT10g) the computation of term frequency (TF) and inverse document frequency (IDF) is not difficult. However, when the corpus is the entire web, direct IDF calculation is impossible and…
Specificity is important for extracting collocations, keyphrases, multi-word and index terms [Newman et al. 2012]. It is also useful for tagging, ontology construction [Ryu and Choi 2006], and automatic summarization of documents [Louis and…