English
Related papers

Related papers: Turkronicles: Diachronic Resources for the Fast Ev…

200 papers

Understanding procedural natural language (e.g., step-by-step instructions) is a crucial step to execution and planning. However, while there are ample corpora and downstream tasks available in English, the field lacks such resources for…

Computation and Language · Computer Science 2024-03-08 Arda Uzunoglu , Gözde Gül Şahin

Large, diachronic datasets of political discourse are hard to come across, especially for resource-lean languages such as Greek. In this paper, we introduce a curated dataset of the Greek Parliament Proceedings that extends chronologically…

Computation and Language · Computer Science 2022-10-25 Konstantina Dritsa , Kaiti Thoma , John Pavlopoulos , Panos Louridas

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

Computation and Language · Computer Science 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

Digital Libraries · Computer Science 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to…

Computation and Language · Computer Science 2018-09-12 Yonatan Belinkov , Alexander Magidow , Alberto Barrón-Cedeño , Avi Shmidman , Maxim Romanov

The availability of large diachronic corpora has provided the impetus for a growing body of quantitative research on language evolution and meaning change. The central quantities in this research are token frequencies of linguistic elements…

Computation and Language · Computer Science 2020-06-17 Andres Karjus , Richard A. Blythe , Simon Kirby , Kenny Smith

It is commonly acknowledged that temporal expression extractors are important components of larger natural language processing systems like information retrieval and question answering systems. Extraction and normalization of temporal…

Computation and Language · Computer Science 2015-09-04 Dilek Küçük , Doğan Küçük

In this study, we develop and assess new corpus selection and training methodologies to improve the effectiveness of Turkish language models. Specifically, we adapted Large Language Model generated datasets and translated English datasets…

Computation and Language · Computer Science 2024-12-05 H. Toprak Kesgin , M. Kaan Yuce , Eren Dogan , M. Egemen Uzun , Atahan Uz , Elif Ince , Yusuf Erdem , Osama Shbib , Ahmed Zeer , M. Fatih Amasyali

Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the under-researched class. In this paper, we address the problem of the absence of organized benchmarks in the Turkish…

Computation and Language · Computer Science 2022-03-17 Ali Safaya , Emirhan Kurtuluş , Arda Göktoğan , Deniz Yuret

We present an expanded version of our previously released Kazakh text-to-speech (KazakhTTS) synthesis corpus. In the new KazakhTTS2 corpus, the overall size has increased from 93 hours to 271 hours, the number of speakers has risen from two…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Saida Mussakhojayeva , Yerbolat Khassanov , Huseyin Atakan Varol

Recent advances in neural machine translation (NMT) have pushed the quality of machine translation systems to the point where they are becoming widely adopted to build competitive systems. However, there is still a large number of languages…

Increased popularity of different text representations has also brought many improvements in Natural Language Processing (NLP) tasks. Without need of supervised data, embeddings trained on large corpora provide us meaningful relations to be…

Computation and Language · Computer Science 2020-02-14 Gökhan Güler , A. Cüneyd Tantuğ

Turkic languages exhibit extensive and diverse etymological relationships among lexical items. These relationships make the Turkic languages promising for exploring automated translation lexicon induction by leveraging cognate and other…

Computation and Language · Computer Science 2015-01-15 Benjamin S. Mericli , Michael Bloodgood

In this study, we present TURead, an eye movement dataset of silent and oral sentence reading in Turkish, an agglutinative language with a shallow orthography understudied in reading research. TURead provides empirical data to investigate…

Computation and Language · Computer Science 2023-12-05 Cengiz Acarturk , Aysegul Ozkan , Tugce Nur Pekcetin , Zuhal Ormanoglu , Bilal Kirkici

This study is an attempt to build a contemporary linguistic corpus for Arabic language. The corpus produced, is a text corpus includes more than five million newspaper articles. It contains over a billion and a half words in total, out of…

Computation and Language · Computer Science 2016-11-15 Ibrahim Abu El-khair

All natural language processing systems (such as parsers, generators, taggers) need to have access to a lexicon about the words in the language. This thesis presents a lexicon architecture for natural language processing in Turkish. Given a…

cmp-lg · Computer Science 2008-02-03 Abdullah Kurtulus Yorulmaz

Measuring how semantics of words change over time improves our understanding of how cultures and perspectives change. Diachronic word embeddings help us quantify this shift, although previous studies leveraged substantial temporally…

Computation and Language · Computer Science 2025-06-17 David Dukić , Ana Barić , Marko Čuljak , Josip Jukić , Martin Tutek

Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fidelity. Prior studies…

Computation and Language · Computer Science 2026-02-09 Duygu Altinok

Diachronic word embeddings -- vector representations of words over time -- offer remarkable insights into the evolution of language and provide a tool for quantifying sociocultural change from text documents. Prior work has used such…

Computation and Language · Computer Science 2020-10-05 Sandeep Soni , Kristina Lerman , Jacob Eisenstein

Our study utilizes deep learning methods for the automated transcription of late nineteenth- and early twentieth-century periodicals written in Arabic script Ottoman Turkish (OT) using the Transkribus platform. We discuss the historical…

Computation and Language · Computer Science 2020-11-03 Suphan Kirmizialtin , David Wrisley