中文
相关论文

相关论文: Mapping Topic Evolution Across Poetic Traditions

200 篇论文

Language change is influenced by many factors, but often starts from synchronic variation, where multiple linguistic patterns or forms coexist, or where different speech communities use language in increasingly different ways. Besides…

社会与信息网络 · 计算机科学 2023-09-06 Andres Karjus , Christine Cuskley

Dialects represent a significant component of human culture and are found across all regions of the world. In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022). However, despite cultural importance,…

计算与语言 · 计算机科学 2025-09-18 Minh Duc Bui , Carolin Holtermann , Valentin Hofmann , Anne Lauscher , Katharina von der Wense

This paper proposes a nonparametric Bayesian method for exploratory data analysis and feature construction in continuous time series. Our method focuses on understanding shared features in a set of time series that exhibit significant…

机器学习 · 统计学 2010-08-13 Suchi Saria , Daphne Koller , Anna Penn

We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000…

计算与语言 · 计算机科学 2019-08-21 Peter D. Turney , Saif M. Mohammad

Grammatical forms are said to evolve via two main mechanisms. These are, respectively, the `descent' mechanism, where current forms can be seen to have descended (albeit with occasional modifications) from their roots in ancient languages,…

统计力学 · 物理学 2023-02-20 Jean-Marc Luck , Anita Mehta

The study uses the British National Corpus 2014, a large sample of contemporary spoken British English, to investigate language patterns across different age groups. Our research attempts to explore how language patterns vary between…

计算与语言 · 计算机科学 2025-06-24 MingZe Tang

We propose a parsimonious topic model for text corpora. In related models such as Latent Dirichlet Allocation (LDA), all words are modeled topic-specifically, even though many words occur with similar frequencies across different topics.…

机器学习 · 计算机科学 2016-05-16 Hossein Soleimani , David J. Miller

Hedges are widely studied across registers and disciplines, yet research on the translation of hedges in political texts is extremely limited. This contrastive study is dedicated to investigating whether there is a diachronic change in the…

计算与语言 · 计算机科学 2023-05-24 Zhaokun Jiang , Ziyin Zhang

Word clouds became a standard tool for presenting results of natural language processing methods such as topic modelling. They exhibit most important words, where word size is often chosen proportional to the relevance of words within a…

统计计算 · 统计学 2023-02-14 Peter Winker

This paper employs two major natural language processing techniques, topic modeling and clustering, to find patterns in folktales and reveal cultural relationships between regions. In particular, we used Latent Dirichlet Allocation and…

计算与语言 · 计算机科学 2022-06-10 Jacob Werzinsky , Zhiyan Zhong , Xuedan Zou

Strategic diagrams and co-word analysis are widely employed to examine the conceptual structure of scientific domains and their development over time. Yet a structural inconsistency characterises dominant longitudinal implementations:…

社会与信息网络 · 计算机科学 2026-03-09 Massimo Aria , Luca D'Aniello , Michelangelo Misuraca , Maria Spano

The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements…

计算与语言 · 计算机科学 2020-06-17 Andres Karjus , Richard A. Blythe , Simon Kirby , Kenny Smith

Latent Dirichlet Allocation (LDA) model is a famous model in the topic model field, it has been studied for years due to its extensive application value in industry and academia. However, the mathematical derivation of LDA model is…

信息检索 · 计算机科学 2019-08-28 Chen Ma

Academic researchers often need to face with a large collection of research papers in the literature. This problem may be even worse for postgraduate students who are new to a field and may not know where to start. To address this problem,…

计算与语言 · 计算机科学 2016-09-30 Leonard K. M. Poon , Nevin L. Zhang

Estimating the semantic similarity between text data is one of the challenging and open research problems in the field of Natural Language Processing (NLP). The versatility of natural language makes it difficult to define rule-based methods…

计算与语言 · 计算机科学 2021-02-24 Dhivya Chandrasekaran , Vijay Mago

Across languages, multiple consecutive adjectives modifying a noun (e.g. "the big red dog") follow certain unmarked ordering rules. While explanatory accounts have been put forward, much of the work done in this area has relied primarily on…

计算与语言 · 计算机科学 2020-10-13 Jun Yen Leung , Guy Emerson , Ryan Cotterell

By determining which were the most common English words and phrases since the beginning of the 16th century, we obtain a unique large-scale view of the evolution of written text. We find that the most common words and phrases in any given…

物理与社会 · 物理学 2012-12-10 Matjaz Perc

The impact of text length on the estimation of lexical diversity has captured the attention of the scientific community for more than a century. Numerous indices have been proposed, and many studies have been conducted to evaluate them, but…

计算与语言 · 计算机科学 2023-08-01 Yves Bestgen

Latent Dirichlet allocation (LDA) is an important hierarchical Bayesian model for probabilistic topic modeling, which attracts worldwide interests and touches on many important applications in text mining, computer vision and computational…

机器学习 · 计算机科学 2015-03-19 Jia Zeng , William K. Cheung , Jiming Liu

We summarize our exploratory investigation into whether Machine Learning (ML) techniques applied to publicly available professional text can substantially augment strategic planning for astronomy. We find that an approach based on Latent…

数字图书馆 · 计算机科学 2024-07-04 Brian Thomas , Harley Thronson , Anthony Buonomo , Louis Barbier