English
Related papers

Related papers: Variable-Length Codes Independent or Closed with r…

200 papers

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith

The impact of text length on the estimation of lexical diversity has captured the attention of the scientific community for more than a century. Numerous indices have been proposed, and many studies have been conducted to evaluate them, but…

Computation and Language · Computer Science 2023-08-01 Yves Bestgen

This paper studies reliability-guaranteed decoding for variable-length stop-feedback (VLSF) codes over correlated noncoherent fading channels. The decoding rule is based on the evolution of the information density associated with a given…

Information Theory · Computer Science 2026-04-20 Guodong Sun , Samir M. Perlaza , Philippe Mary , Jean-Marie Gorce

We consider strategies to organize easily updatable associative arrays in external memory. These arrays are used for full-text search. We study indexes with different keys: single word form, two word forms, and sequences of word forms. The…

Information Retrieval · Computer Science 2020-07-21 Alexander B. Veretennikov

Given a finite alphabet $A$, a quasi-metric $d$ over $A^*$, and a non-negative integer $k$, we introduce the relation $\tau_{d,k}\subseteq A^*\times A^*$ such that $(x,y)\in\tau_{d,k}$ holds whenever $d(x,y)\le k$. The error detection…

Discrete Mathematics · Computer Science 2023-09-06 Jean Néraud

Batch codes are of potential use for load balancing and private information retrieval in distributed data storage systems. Recently, a special case of batch codes, termed functional batch codes, was proposed in the literature. In functional…

Information Theory · Computer Science 2026-05-26 Kristiina Oksner , Henk D. L. Hollmann , Ago-Erik Riet , Vitaly Skachek

Cut-set bounds on achievable rates for network communication protocols are not in general tight. In this paper we introduce a new technique for proving converses for the problem of transmission of correlated sources in networks, that…

Information Theory · Computer Science 2011-05-31 Amin Aminzadeh Gohari , Shenghao Yang , Sidharth Jaggi

We consider the problem of bounding large deviations for non-i.i.d. random variables that are allowed to have arbitrary dependencies. Previous works typically assumed a specific dependence structure, namely the existence of independent…

Probability · Mathematics 2018-11-06 Christoph H. Lampert , Liva Ralaivola , Alexander Zimin

Statistical language modeling techniques have successfully been applied to source code, yielding a variety of new software development tools, such as tools for code suggestion and improving readability. A major issue with these techniques…

Software Engineering · Computer Science 2019-03-15 Rafael-Michael Karampatsis , Charles Sutton

The number of random bits required to approximate a target distribution in terms of un-normalized informational divergence is considered. It is shown that for a variable-to-variable length encoder, this number is lower bounded by the…

Information Theory · Computer Science 2013-08-02 Georg Böcherer , Rana Ali Amjad

A self-dual binary linear code is called Type I code if it has singly-even codewords, i.e.~it has codewords with weight divisible by $2.$ The purpose of this paper is to investigate interesting properties of Type I codes of different…

Information Theory · Computer Science 2021-10-19 Carolin Hannusch , Roland S. Major

Qualitative data analysis provides insight into the underlying perceptions and experiences within unstructured data. However, the time-consuming nature of the coding process, especially for larger datasets, calls for innovative approaches,…

Human-Computer Interaction · Computer Science 2024-03-12 Elisabeth Kirsten , Annalina Buckmann , Abraham Mhaidli , Steffen Becker

The zero-error channel capacity is the maximum asymptotic rate that can be reached with error probability exactly zero, instead of a vanishing error probability. The nature of this problem, essentially combinatorial rather than…

Information Theory · Computer Science 2020-01-13 Nicolas Charpenay , Maël Le Treust

Degraded text recognition is a difficult task. Given a noisy text image, a word recognizer can be applied to generate several candidates for each word image. High-level knowledge sources can then be used to select a decision from the…

cmp-lg · Computer Science 2008-02-03 Tao Hong

Recently, it has been claimed that a linear relationship between a measure of information content and word length is expected from word length optimization and it has been shown that this linearity is supported by a strong correlation…

Data Analysis, Statistics and Probability · Physics 2019-12-11 Ramon Ferrer-i-Cancho , Fermín Moscoso del Prado Martín

We consider the question of interactive communication, in which two remote parties perform a computation while their communication channel is (adversarially) noisy. We extend here the discussion into a more general and stronger class of…

Data Structures and Algorithms · Computer Science 2016-05-25 Mark Braverman , Ran Gelles , Jieming Mao , Rafail Ostrovsky

This study investigates the fundamental limits of variable-length compression in which prefix-free constraints are not imposed (i.e., one-to-one codes are studied) and non-vanishing error probabilities are permitted. Due in part to a…

Information Theory · Computer Science 2021-10-05 Yuta Sakai , Recep Can Yavas , Vincent Y. F. Tan

We address the recently suggested problem of causal lossless coding of a randomly arriving source samples. We construct variable-to-fixed coding schemes and show that they outperform the previously considered fixed-to-variable schemes when…

Information Theory · Computer Science 2020-10-27 Uri Abend , Anatoly Khina

Lossless variable-length source coding with unequal cost function is considered for general sources. In this problem, the codeword cost instead of codeword length is important. The infimum of average codeword cost has already been…

Information Theory · Computer Science 2012-05-10 Ryo Nomura , Toshiyasu Matsushima

Large language models (LLMs) are being increasingly adopted in the software engineering domain, yet the robustness of their grasp on core software design concepts remains unclear. We conduct an empirical study to systematically evaluate…

Software Engineering · Computer Science 2025-12-30 Mootez Saad , Boqi Chen , José Antonio Hernández López , Dániel Varró , Tushar Sharma