English
Related papers

Related papers: Diversity, Density, and Homogeneity: Quantitative …

200 papers

Recent work on vector-based compositional natural language semantics has proposed the use of density matrices to model lexical ambiguity and (graded) entailment (e.g. Piedeleu et al 2015, Bankova et al 2019, Sadrzadeh et al 2018). Ambiguous…

Computation and Language · Computer Science 2020-11-06 Adriana D. Correia , Michael Moortgat , Henk T. C. Stoof

Discourse cohesion facilitates text comprehension and helps the reader form a coherent narrative. In this study, we aim to computationally analyze the discourse cohesion in scientific scholarly texts using multilayer network representation…

Computation and Language · Computer Science 2022-11-09 Vasudha Bhatnagar , Swagata Duari , S. K. Gupta

Various measures of dispersion have been proposed to paint a fuller picture of a word's distribution in a corpus, but only little has been done to validate them externally. We evaluate a wide range of dispersion measures as predictors of…

Computation and Language · Computer Science 2025-01-14 Adam Nohejl , Taro Watanabe

Large pre-trained language models (LMs) have demonstrated impressive capabilities in generating long, fluent text; however, there is little to no analysis on their ability to maintain entity coherence and consistency. In this work, we focus…

Computation and Language · Computer Science 2022-02-04 Pinelopi Papalampidi , Kris Cao , Tomas Kocisky

Text style transfer involves rewriting the content of a source sentence in a target style. Despite there being a number of style tasks with available data, there has been limited systematic discussion of how text style datasets relate to…

Computation and Language · Computer Science 2021-08-19 Stephanie Schoch , Wanyu Du , Yangfeng Ji

Large language models are increasingly being used to label or rate psychological features in text data. This approach helps address one of the limiting factors of digital trace data - their lack of an inherent target of measurement.…

Human-Computer Interaction · Computer Science 2024-10-15 Joseph J. P. Simons , Wong Liang Ze , Prasanta Bhattacharya , Brandon Siyuan Loh , Wei Gao

Matching for causal inference is a well-studied problem, but standard methods fail when the units to match are text documents: the high-dimensional and rich nature of the data renders exact matching infeasible, causes propensity scores to…

Methodology · Statistics 2019-03-15 Reagan Mozer , Luke Miratrix , Aaron Russell Kaufman , L. Jason Anastasopoulos

Text classification is one of the most widely studied tasks in natural language processing. Motivated by the principle of compositionality, large multilayer neural network models have been employed for this task in an attempt to effectively…

Computation and Language · Computer Science 2018-08-07 Devendra Singh Sachan , Manzil Zaheer , Ruslan Salakhutdinov

Recent works have proposed that activations in language models can be modelled as sparse linear combinations of vectors corresponding to features of input text. Under this assumption, these works aimed to reconstruct feature directions…

Machine Learning · Computer Science 2023-10-17 Mingyang Deng , Lucas Tao , Joe Benton

We investigate how quantum coherence can be distributed among the several off-diagonal elements of an arbitrary density matrix. An easily computable quantity that captures this variability notion is proposed and it is argued that it…

Quantum Physics · Physics 2026-04-24 Fernando Parisio

This paper introduces a new statistical approach to partitioning text automatically into coherent segments. Our approach enlists both short-range and long-range language models to help it sniff out likely sites of topic changes in text. To…

cmp-lg · Computer Science 2008-02-03 Doug Beeferman , Adam Berger , John Lafferty

Current trends in pre-training Large Language Models (LLMs) primarily focus on the scaling of model and dataset size. While the quality of pre-training data is considered an important factor for training powerful LLMs, it remains a nebulous…

Computation and Language · Computer Science 2025-07-04 Brando Miranda , Alycia Lee , Sudharsan Sundar , Allison Casasola , Rylan Schaeffer , Elyas Obbad , Sanmi Koyejo

This paper considers the problem of specifying a simple approximating density function for a given data set (x_1,...,x_n). Simplicity is measured by the number of modes but several different definitions of approximation are introduced. The…

Statistics Theory · Mathematics 2007-06-13 P. Laurie Davies , Arne Kovac

Diversities have recently been developed as multiway metrics admitting clear and useful notions of hyperconvexity and tight span. In this note we consider the analytic properties of diversities, in particular the generalizations of uniform…

Metric Geometry · Mathematics 2013-11-19 Andrew Poelstra

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

In this paper we exploit concepts of information theory to address the fundamental problem of identifying and defining the most suitable tools to extract, in a automatic and agnostic way, information from a generic string of characters. We…

Statistical Mechanics · Physics 2009-11-10 Andrea Baronchelli , Emanuele Caglioti , Vittorio Loreto

As an intrinsic and fundamental property of big data, data heterogeneity exists in a variety of real-world applications, such as precision medicine, autonomous driving, financial applications, etc. For machine learning algorithms, the…

Machine Learning · Computer Science 2023-04-04 Jiashuo Liu , Jiayun Wu , Bo Li , Peng Cui

A stereotype is a generalized perception of a specific group of humans. It is often potentially encoded in human language, which is more common in texts on social issues. Previous works simply define a sentence as stereotypical and…

Computation and Language · Computer Science 2024-01-30 Yang Liu

Automatically evaluating the coherence of summaries is of great significance both to enable cost-efficient summarizer evaluation and as a tool for improving coherence by selecting high-scoring candidate summaries. While many different…

Computation and Language · Computer Science 2022-09-16 Julius Steen , Katja Markert

Understanding the role of demographic diversity in group settings requires effective quantitative metrics. Intersectional feminist theory has highlighted that demographic identities can intersect in complex ways, but most metrics used to…

General Mathematics · Mathematics 2025-09-19 Leah Hoogstra , Katherine Slyman , Bjorn Sandstede
‹ Prev 1 3 4 5 6 7 10 Next ›