English
Related papers

Related papers: Turkronicles: Diachronic Resources for the Fast Ev…

200 papers

Neural information retrieval systems excel in high-resource languages but remain underexplored for morphologically rich, lower-resource languages such as Turkish. Dense bi-encoders currently dominate Turkish IR, yet late-interaction models…

Computation and Language · Computer Science 2025-11-21 Özay Ezerceli , Mahmoud El Hussieni , Selva Taş , Reyhan Bayraktar , Fatma Betül Terzioğlu , Yusuf Çelebi , Yağız Asker

Recent advances in natural language processing (NLP) have increasingly enabled LegalTech applications, yet existing studies specific to Turkish law have still been limited due to the scarcity of domain-specific data and models. Although…

Computation and Language · Computer Science 2026-04-07 Mehmet Utku Öztürk , Tansu Türkoğlu , Buse Buz-Yalug

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

It is offered to consider word meanings changes in diachrony as semicontinuous random walk with reflecting and swallowing screens. The basic characteristics of word life cycle are defined. Verification of the model has been realized on the…

Computation and Language · Computer Science 2007-05-23 Victor Kromer

This paper provides a global vision of the scientific publications related with the Systemic Lupus Erythematosus (SLE), taking as starting point abstracts of articles. Through the time, abstracts have been evolving towards higher complexity…

Computation and Language · Computer Science 2016-07-27 Daria Micaela Hernandez , Monica Becue-Bertaut , Igor Barahona

Most Arabic natural language processing tools and resources are developed to serve Modern Standard Arabic (MSA), which is the official written language in the Arab World. Some Dialectal Arabic varieties, notably Egyptian Arabic, have…

Computation and Language · Computer Science 2016-09-13 Salam Khalifa , Nizar Habash , Dana Abdulrahim , Sara Hassan

The paper introduces manually annotated test sets for the task of tracing diachronic (temporal) semantic shifts in Russian. The two test sets are complementary in that the first one covers comparatively strong semantic changes occurring to…

Computation and Language · Computer Science 2019-07-31 Vadim Fomin , Daria Bakshandaeva , Julia Rodina , Andrey Kutuzov

The evolution of Turkish Journal of Chemistry (Turk J. Chem) Hirsch index (h-index) over the period 1995-2005 is studied and determined in the case of the self and without self-citations. It is seen that the effect of Hirsch index of Turk…

Physics Education · Physics 2007-05-23 Metin Orbay , Orhan Karamustafaoglu , Feda Oner

In this thesis, we developed a comprehensive framework for sentiment analysis that takes its many aspects into account mainly for Turkish. We have also proposed several approaches specific to sentiment analysis in English only. We have…

Computation and Language · Computer Science 2025-12-02 Cem Rifki Aydin

We analyze the occurrence frequencies of over 15 million words recorded in millions of books published during the past two centuries in seven different languages. For all languages and chronological subsets of the data we confirm that two…

Physics and Society · Physics 2012-12-12 Alexander M. Petersen , Joel N. Tenenbaum , Shlomo Havlin , H. Eugene Stanley , Matjaz Perc

All living languages change over time. The causes for this are many, one being the emergence and borrowing of new linguistic elements. Competition between the new elements and older ones with a similar semantic or grammatical function may…

Computation and Language · Computer Science 2020-06-17 Andres Karjus , Richard A. Blythe , Simon Kirby , Kenny Smith

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

Computation and Language · Computer Science 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

Hedges are widely studied across registers and disciplines, yet research on the translation of hedges in political texts is extremely limited. This contrastive study is dedicated to investigating whether there is a diachronic change in the…

Computation and Language · Computer Science 2023-05-24 Zhaokun Jiang , Ziyin Zhang

A key challenge when trying to understand innovation is that it is a dynamic, ongoing process, which can be highly contingent on ephemeral factors such as culture, economics, or luck. This means that any analysis of the real-world process…

Artificial Intelligence · Computer Science 2023-10-16 Paul Kinsler

Has the style of scientific communication changed due to the growing use of large language models in the writing process? We address this question in the domain of Natural Language Processing by leveraging two data resources we create: a…

Computation and Language · Computer Science 2026-05-20 Filip Miletić , Neele Falk

We present a dataset that contains every instance of all tokens (~ words) ever written in undeleted, non-redirect English Wikipedia articles until October 2016, in total 13,545,349,787 instances. Each token is annotated with (i) the article…

Computation and Language · Computer Science 2017-03-27 Fabian Flöck , Kenan Erdogan , Maribel Acosta

We present a hybrid methodology for generating large-scale semantic relationship datasets in low-resource languages, demonstrated through a comprehensive Turkish semantic relations corpus. Our approach integrates three phases: (1) FastText…

Computation and Language · Computer Science 2026-01-21 Ebubekir Tosun , Mehmet Emin Buldur , Özay Ezerceli , Mahmoud ElHussieni

Languages emerge and change over time at the population level though interactions between individual speakers. It is, however, hard to directly observe how a single speaker's linguistic innovation precipitates a population-wide change in…

Computation and Language · Computer Science 2021-06-04 Richard A Blythe , William Croft

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent resources: a corpus…

Computation and Language · Computer Science 2022-07-12 Elisa Gugliotta , Marco Dinarelli

Syntactic discontinuity is a grammatical phenomenon in which a constituent is split into more than one part because of the insertion of an element which is not part of the constituent. This is observed in many languages across the world…

Computation and Language · Computer Science 2025-06-06 Ratna Kandala , Prakash Mondal