English
Related papers

Related papers: The Danish Gigaword Project

200 papers

Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown…

Computation and Language · Computer Science 2021-10-14 Damián Blasi , Antonios Anastasopoulos , Graham Neubig

This paper investigates the impact of corpus creation decisions on large multi-lingual geographic web corpora. Beginning with a 427 billion word corpus derived from the Common Crawl, three methods are used to improve the quality of…

Computation and Language · Computer Science 2024-03-14 Jonathan Dunn

Learning high-quality domain word embeddings is important for achieving good performance in many NLP tasks. General-purpose embeddings trained on large-scale corpora are often sub-optimal for domain-specific applications. However,…

Computation and Language · Computer Science 2018-05-28 Hu Xu , Bing Liu , Lei Shu , Philip S. Yu

Named entity recognition is one of the cornerstones of Danish NLP, essential for language technology applications within both industry and research. However, Danish NER is inhibited by a lack of available datasets. As a consequence, no…

Computation and Language · Computer Science 2024-02-29 Kenneth Enevoldsen , Emil Trenckner Jessen , Rebekah Baglini

As language data and associated technologies proliferate and as the language resources community expands, it is becoming increasingly difficult to locate and reuse existing resources. Are there any lexical resources for such-and-such a…

Computation and Language · Computer Science 2007-05-23 Steven Bird , Gary Simons

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

Computation and Language · Computer Science 2020-05-04 Anoop Kunchukuttan , Divyanshu Kakwani , Satish Golla , Gokul N. C. , Avik Bhattacharyya , Mitesh M. Khapra , Pratyush Kumar

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Despite the rapid expansion of types of large language models, there remains a notable gap in models specifically designed for the Dutch language. This gap is not only a shortage in terms of pretrained Dutch models but also in terms of…

Computation and Language · Computer Science 2023-12-21 Bram Vanroy

Large text corpora are increasingly important for a wide variety of Natural Language Processing (NLP) tasks, and automatic language identification (LangID) is a core technology needed to collect such datasets in a multilingual context.…

Computation and Language · Computer Science 2020-10-30 Isaac Caswell , Theresa Breiner , Daan van Esch , Ankur Bapna

This paper presents the "Leipzig Corpus Miner", a technical infrastructure for supporting qualitative and quantitative content analysis. The infrastructure aims at the integration of 'close reading' procedures on individual documents with…

Computation and Language · Computer Science 2017-07-12 Andreas Niekler , Gregor Wiedemann , Gerhard Heyer

Discourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations. This work investigates whether large language models (LLMs) capture discourse knowledge…

Computation and Language · Computer Science 2025-06-05 Florian Eichin , Yang Janet Liu , Barbara Plank , Michael A. Hedderich

Parliamentary debates represent a large and partly unexploited treasure trove of publicly accessible texts. In the German-speaking area, there is a certain deficit of uniformly accessible and annotated corpora covering all German-speaking…

Computation and Language · Computer Science 2022-04-25 Giuseppe Abrami , Mevlüt Bagci , Leon Hammerla , Alexander Mehler

Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack sufficient diachronic information. We develop a large-scale,…

Computation and Language · Computer Science 2016-12-30 Yonatan Belinkov , Alexander Magidow , Maxim Romanov , Avi Shmidman , Moshe Koppel

The word embedding methods have been proven to be very useful in many tasks of NLP (Natural Language Processing). Much has been investigated about word embeddings of English words and phrases, but only little attention has been dedicated to…

Computation and Language · Computer Science 2016-08-03 Lukáš Svoboda , Tomáš Brychcín

We present SnakModel, a Danish large language model (LLM) based on Llama2-7B, which we continuously pre-train on 13.6B Danish words, and further tune on 3.7M Danish instructions. As best practices for creating LLMs for smaller language…

Computation and Language · Computer Science 2024-12-18 Mike Zhang , Max Müller-Eberstein , Elisa Bassignana , Rob van der Goot

The need for raw large raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to…

Computation and Language · Computer Science 2022-01-19 Julien Abadji , Pedro Ortiz Suarez , Laurent Romary , Benoît Sagot

Deep language models have achieved remarkable success in the NLP domain. The standard way to train a deep language model is to employ unsupervised learning from scratch on a large unlabeled corpus. However, such large corpora are only…

Computation and Language · Computer Science 2021-12-02 Arda Akdemir , Yeojoo Jeon

Word meaning is notoriously difficult to capture, both synchronically and diachronically. In this paper, we describe the creation of the largest resource of graded contextualized, diachronic word meaning annotation in four different…

Computation and Language · Computer Science 2024-07-09 Dominik Schlechtweg , Nina Tahmasebi , Simon Hengchen , Haim Dubossarsky , Barbara McGillivray

Recent natural language processing (NLP) techniques have accomplished high performance on benchmark datasets, primarily due to the significant improvement in the performance of deep learning. The advances in the research community have led…

Computation and Language · Computer Science 2022-10-24 Marwan Omar , Soohyeon Choi , DaeHun Nyang , David Mohaisen

Many European languages possess rich biblical translation histories, yet existing corpora - in prioritizing linguistic breadth - often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament…

Computation and Language · Computer Science 2026-05-14 Maciej Rapacz , Aleksander Smywiński-Pohl