English
Related papers

Related papers: Quantifying literature quality using complexity cr…

200 papers

This article addresses Second Language (L2) writing development through an investigation of new grammatical and structural complexity metrics. We explore the paradigmatic production in learner English by linking language functions to…

Computation and Language · Computer Science 2025-03-17 Cyriel Mallart , Andrew Simpkin , Nicolas Ballier , Paula Lissón , Rémi Venant , Jen-Yu Li , Bernardo Stearns , Thomas Gaillat

In this paper, we use topological data analysis (TDA) tools such as persistent homology, persistent entropy and bottleneck distance, to provide a {\it TDA-based summary} of any given set of texts and a general method for computing a…

Computation and Language · Computer Science 2022-03-15 Eduardo Paluzo-Hidalgo , Rocio Gonzalez-Diaz , Miguel A. Gutiérrez-Naranjo

We analyze the concept of virtuosity as a collective attribute in music and its relationship with the entropy based on an experiment that compares two sets of digital signals played by composer-performer electric guitarists. Based on an…

Sound · Computer Science 2024-04-26 Igor Lugo , Martha G. Alatriste-Contreras

Contextual entropy is a psycholinguistic measure capturing the anticipated difficulty of processing a word just before it is encountered. Recent studies have tested for entropy-related effects as a potential complement to well-known effects…

Computation and Language · Computer Science 2025-07-31 Christian Clark , Byung-Doh Oh , William Schuler

A citation-based indicator for interdisciplinarity has been missing hitherto among the set of available journal indicators. In this study, we investigate network indicators (betweenness centrality), journal indicators (Shannon entropy, the…

Digital Libraries · Computer Science 2010-09-22 Loet Leydesdorff , Ismael Rafols

Text summarization has a wide range of applications in many scenarios. The evaluation of the quality of the generated text is a complex problem. A big challenge to language evaluation is that there is a clear divergence between existing…

Computation and Language · Computer Science 2023-09-20 Ning Wu , Ming Gong , Linjun Shou , Shining Liang , Daxin Jiang

Scholars, awards committees, and laypeople frequently discuss the merit of written works. Literary professionals and journalists differ in how much perspectivism they concede in their book reviews. Here, we quantify how strongly book…

Digital Libraries · Computer Science 2025-03-05 Hannes Rosenbusch , Luke Korthals

How can we measure whether a natural language generation system produces both high quality and diverse outputs? Human evaluation captures quality but not diversity, as it does not catch models that simply plagiarize from the training set.…

Computation and Language · Computer Science 2019-04-08 Tatsunori B. Hashimoto , Hugh Zhang , Percy Liang

Recently, amounts of works utilize perplexity~(PPL) to evaluate the quality of the generated text. They suppose that if the value of PPL is smaller, the quality(i.e. fluency) of the text to be evaluated is better. However, we find that the…

Computation and Language · Computer Science 2023-03-16 Yequan Wang , Jiawen Deng , Aixin Sun , Xuying Meng

The use of naive Bayesian classifier (NB) and the classifier by the k nearest neighbors (kNN) in classification semantic analysis of authors' texts of English fiction has been analysed. The authors' works are considered in the vector space…

Computation and Language · Computer Science 2012-10-23 Bohdan Pavlyshenko

We present in this paper a numerical investigation of literary texts by various well-known English writers, covering the first half of the twentieth century, based upon the results obtained through corpus analysis of the texts. A fractal…

Other Condensed Matter · Physics 2009-11-11 L. L. Goncalves , L. B. Goncalves

Scientific literature review generation aims to extract and organize important information from an abundant collection of reference papers and produces corresponding reviews while lacking a clear and logical hierarchy. We observe that a…

Computation and Language · Computer Science 2023-11-20 Kun Zhu , Xiaocheng Feng , Xiachong Feng , Yingsheng Wu , Bing Qin

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith

A lot of scientific works are published in different areas of science, technology, engineering and mathematics. It is not easy, even for experts, to judge the quality of authors, papers and venues (conferences and journals). An objective…

Digital Libraries · Computer Science 2017-08-04 Arindam Pal , Sushmita Ruj

We examine, analyze, and compare four representative creativity measures--perplexity, LLM-as-a-Judge, the Creativity Index (CI; measuring n-gram overlap with web corpora), and syntactic templates (detecting repetition of common…

Computation and Language · Computer Science 2026-01-29 Li-Chun Lu , Miri Liu , Pin-Chun Lu , Yufei Tian , Shao-Hua Sun , Nanyun Peng

We use large language models (LLMs) to uncover long-ranged structure in English texts from a variety of sources. The conditional entropy or code length in many cases continues to decrease with context length at least to $N\sim 10^4$…

Statistical Mechanics · Physics 2026-01-01 Colin Scheibner , Lindsay M. Smith , William Bialek

The review summarizes the main methodological concepts used in studying natural language from the perspective of complexity science and documents their applicability in identifying both universal and system-specific features of language in…

Physics and Society · Physics 2024-01-09 Tomasz Stanisz , Stanisław Drożdż , Jarosław Kwapień

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

Koch and Oesterreicher's model of "N\"ahe und Distanz" (N\"ahe = immediacy, conceptual orality; Distanz = distance, conceptual literacy) is constantly used in German linguistics. However, there is no statistical foundation for use in corpus…

Computation and Language · Computer Science 2025-02-06 Volker Emmrich

Theoretical work in morphological typology offers the possibility of measuring morphological diversity on a continuous scale. However, literature in Natural Language Processing (NLP) typically labels a whole language with a strict type of…

Computation and Language · Computer Science 2022-05-09 Arturo Oncevay , Duygu Ataman , Niels van Berkel , Barry Haddow , Alexandra Birch , Johannes Bjerva
‹ Prev 1 8 9 10 Next ›