English
Related papers

Related papers: Spectral Analysis of Word Statistics

200 papers

The words-as-classifiers model of grounded lexical semantics learns a semantic fitness score between physical entities and the words that are used to denote those entities. In this paper, we explore how such a model can incrementally…

Computation and Language · Computer Science 2019-11-11 Daniele Moro , Stacy Black , Casey Kennington

Subwords are the most widely used output units in end-to-end speech recognition. They combine the best of two worlds by modeling the majority of frequent words directly and at the same time allow open vocabulary speech recognition by…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Egor Lakomkin , Jahn Heymann , Ilya Sklyar , Simon Wiesler

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more communities. But word…

Computation and Language · Computer Science 2023-02-14 Thyge Enggaard , August Lohse , Morten Axel Pedersen , Sune Lehmann

A fundamental concept in multivariate statistics, sample correlation matrix, is often used to infer the correlation/dependence structure among random variables, when the population mean and covariance are unknown. A natural block extension…

Statistics Theory · Mathematics 2022-09-09 Zhigang Bao , Jiang Hu , Xiaocong Xu , Xiaozhuo Zhang

We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of…

Computation and Language · Computer Science 2014-11-13 Vivek Kulkarni , Rami Al-Rfou , Bryan Perozzi , Steven Skiena

Theoretical analysis of biological and artificial neural networks e.g. modelling of synaptic or weight matrices necessitate consideration of the generic real-asymmetric matrix ensembles, those with varying order of matrix elements e.g. a…

Disordered Systems and Neural Networks · Physics 2025-09-15 Ratul Dutta , Pragya Shukla

We consider expanding maps such that the unit interval can be represented as a full symbolic shift space with bounded distortion. There are already theorems about the Hausdorff dimension for sets defined by the set of accumulation points…

Dynamical Systems · Mathematics 2009-04-29 David Färm

The analysis of thousands of time series in different languages reveals that word usage presents oscillations with a prevalence of 16-year cycles, mounted on slowly varying trends. These components carry different information: while similar…

Neurons and Cognition · Quantitative Biology 2022-07-20 Alejandro Pardo Pintos , Diego E Shalom , Enzo Tagliazucchi , Gabriel Mindlin , Marcos A Trevisan

Data mining allows the exploration of sequences of phenomena, whereas one usually tends to focus on isolated phenomena or on the relation between two phenomena. It offers invaluable tools for theoretical analyses and exploration of the…

Computation and Language · Computer Science 2007-10-15 Catherine Recanati , Nicoleta Rogovschi , Younès Bennani

Word embeddings are powerful representations that form the foundation of many natural language processing architectures, both in English and in other languages. To gain further insight into word embeddings, we explore their stability (e.g.,…

Computation and Language · Computer Science 2021-09-13 Laura Burdick , Jonathan K. Kummerfeld , Rada Mihalcea

Probabilistic approaches to part-of-speech tagging rely primarily on whole-word statistics about word/tag combinations as well as contextual information. But experience shows about 4 per cent of tokens encountered in test sets are unknown…

Computation and Language · Computer Science 2013-02-28 Greg Adams , Beth Millar , Eric Neufeld , Tim Philip

There have been several efforts to extend distributional semantics beyond individual words, to measure the similarity of word pairs, phrases, and sentences (briefly, tuples; ordered sets of words, contiguous or noncontiguous). One way to…

Machine Learning · Computer Science 2013-10-21 Peter D. Turney

In this article we propose a novel method to estimate the frequency distribution of linguistic variables while controlling for statistical non-independence due to shared ancestry. Unlike previous approaches, our technique uses all available…

Populations and Evolution · Quantitative Biology 2021-03-22 Gerhard Jäger , Johannes Wahle

How does word frequency in pre-training data affect the behavior of similarity metrics in contextualized BERT embeddings? Are there systematic ways in which some word relationships are exaggerated or understated? In this work, we explore…

Computation and Language · Computer Science 2021-04-20 Kaitlyn Zhou , Kawin Ethayarajh , Dan Jurafsky

Integer partitions have fascinated people for centuries, from Ramanujan's groundbreaking congruences to the modern theory of modular forms. This paper investigates the statistical properties of odd unimodal sequences--a natural refinement…

Number Theory · Mathematics 2026-05-11 Bing He , Guanting Liu

Distributional semantic models provide vector representations for words by gathering co-occurrence frequencies from corpora of text. Compositional distributional models extend these from words to phrases and sentences. In categorical…

Computation and Language · Computer Science 2018-10-10 Esma Balkir , Dimitri Kartsaklis , Mehrnoosh Sadrzadeh

Traditional linguistic theories have largely regard language as a formal system composed of rigid rules. However, their failures in processing real language, the recent successes in statistical natural language processing, and the findings…

Computation and Language · Computer Science 2020-12-02 Shuiyuan Yu , Chunshan Xu , Haitao Liu

Random matrix theory is finding an increasing number of applications in the context of information theory and communication systems, especially in studying the properties of complex networks. Such properties include short-term and long-term…

Mathematical Physics · Physics 2015-01-13 Sherif M. Abuelenin , Adel Y. Abul-Magd

In this work we seek clusters of genomic words in human DNA by studying their inter-word lag distributions. Due to the particularly spiked nature of these histograms, a clustering procedure is proposed that first decomposes each…

Applications · Statistics 2021-01-13 Ana Helena Tavares , Jakob Raymaekers , Peter J. Rousseeuw , Paula Brito , Vera Afreixo

The distances between words calculated in word units are studied and compared with the distributions of the Random Matrix Theory (RMT). It is found that the distribution of distance between the same words can be well described by the…

Computation and Language · Computer Science 2021-10-27 Bogdan Łobodziński