Related papers: Comment on "Linguistic Features of Noncoding DNA S…
We consider the response of a memoryless nonlinear device that converts an input signal $\xi(t)$ into an output $\eta(t)$ that only depends on the value of the input at the same time, $t$. For input Gaussian noise with power spectrum…
All known life forms are based upon a hierarchy of interwoven feedback loops, operating over a cascade of space, time and energy scales. Among the most basic loops are those connecting DNA and proteins. For example, in genetic networks, DNA…
In psycholinguistic modeling, surprisal from larger pre-trained language models has been shown to be a poorer predictor of naturalistic human reading times. However, it has been speculated that this may be due to data leakage that caused…
The statistical over-representation of phonological features in the basic vocabulary of languages is often interpreted as reflecting potentially universal sound symbolic patterns. However, most of those results have not been tested…
This is a pedagogical review of the ubiquitous 1/f^\alpha noises. The sections include the representation of 1/f^\alpha noise as a superposition of many relaxation processes; a discussion of the infinitely large fluctuations in the low…
Certifying quantum properties from the probability distributions they induce is an important task for several purposes. While this framework has been largely explored and used for quantum states, its extrapolation to the level of channels…
Signal processing (SP) techniques convert DNA and protein sequences into information that lead to successful drug discovery. One must, however, be aware about the difference between information and entropy1. Eight other physical properties…
Despite the widely reported success of embedding-based machine learning methods on natural language processing tasks, the use of more easily interpreted engineered features remains common in fields such as cognitive impairment (CI)…
Time evolution of the cities and of the languages is considered in terms of multiplicative noise and fragmentation processes; where power law (Pareto-Zipf law) and slightly asymmetric log-normal (Gauss) distribution result for the size…
We must recognize that natural language is a way of information encoding, and it encodes not only the information but also the procedures for how information is processed. To understand natural language, the same as we conceive and design…
Zipf's law is just one out of many universal laws proposed to describe statistical regularities in language. Here we review and critically discuss how these laws can be statistically interpreted, fitted, and tested (falsified). The modern…
Discovering the mechanism underlying the ubiquity of $"1/f^{\alpha}"$ noise has been a long--standing problem. The wide range of systems in which the fluctuations show the implied long--time correlations suggests the existence of some…
Recent experiments indicate a connection between the low- and high-frequency noise affecting superconducting quantum systems. We explore the possibilities that both noises can be produced by one ensemble of microscopic modes, made up, e.g.,…
Semi-structured data formats such as JSON have proved to be useful data models for applications that require flexibility in the format of data stored. However, JSON data often come without the schemas that are typically available with…
Inspired by recent developments in natural language processing, we propose a novel approach to sign language processing based on phonological properties validated by American Sign Language users. By taking advantage of datasets composed of…
We present a theoretical and empirical investigation of the statistical behaviour of the words in a text produced by human language. To this aim, we analyse the word distribution of various texts of Italian language selected from a specific…
Building on the computer science concept of code smells, we initiate the study of law smells, i.e., patterns in legal texts that pose threats to the comprehensibility and maintainability of the law. With five intuitive law smells as running…
Emotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed…
Native language identification (NLI) is the task of training (via supervised machine learning) a classifier that guesses the native language of the author of a text. This task has been extensively researched in the last decade, and the…
A foundational assumption in linguistics holds that the relationship between a word's sound and its meaning is arbitrary. Accumulating evidence from sound symbolism challenges this view, yet no study has systematically mapped the…