English
Related papers

Related papers: Compiling and Processing Historical and Contempora…

200 papers

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

Computation and Language · Computer Science 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

Recent advances in natural language processing have raised expectations for generative models to produce coherent text across diverse language varieties. In the particular case of the Portuguese language, the predominance of Brazilian…

Computation and Language · Computer Science 2025-02-21 Hugo Sousa , Rúben Almeida , Purificação Silvano , Inês Cantante , Ricardo Campos , Alípio Jorge

In this work, we show that the difference in performance of embeddings from differently sourced data for a given language can be due to other factors besides data size. Natural language processing (NLP) tasks usually perform better with…

Computation and Language · Computer Science 2020-11-09 Tosin P. Adewumi , Foteini Liwicki , Marcus Liwicki

We illustrate the use of machine learning techniques to analyze, structure, maintain, and evolve a large online corpus of academic literature. An emerging field of research can be identified as part of an existing corpus, permitting the…

Information Retrieval · Computer Science 2009-11-10 Paul Ginsparg , Paul Houle , Thorsten Joachims , Jae-Hoon Sul

We present the IIT Bombay English-Hindi Parallel Corpus. The corpus is a compilation of parallel corpora previously available in the public domain as well as new parallel corpora we collected. The corpus contains 1.49 million parallel…

Computation and Language · Computer Science 2018-05-22 Anoop Kunchukuttan , Pratik Mehta , Pushpak Bhattacharyya

This report describes the MUDOS-NG summarization system, which applies a set of language-independent and generic methods for generating extractive summaries. The proposed methods are mostly combinations of simple operators on a generic…

Computation and Language · Computer Science 2010-12-10 George Giannakopoulos , George Vouros , Vangelis Karkaletsis

Event extraction is an NLP task that commonly involves identifying the central word (trigger) for an event and its associated arguments in text. ACE-2005 is widely recognised as the standard corpus in this field. While other corpora, like…

Computation and Language · Computer Science 2024-09-02 Luís Filipe Cunha , Purificação Silvano , Ricardo Campos , Alípio Jorge

Parallel corpora play an important role in training machine translation (MT) models, particularly for low-resource languages where high-quality bilingual data is scarce. This review provides a comprehensive overview of available parallel…

Computation and Language · Computer Science 2025-04-23 Rahul Raja , Arpita Vats

We describe an algebra for composing automata which includes both classical and quantum entities and their communications. We illustrate by describing in detail a quantum protocol.

Logic in Computer Science · Computer Science 2009-01-30 L. de Francesco Albasini , N. Sabadini , R. F. C. Walters

This paper presents the "Leipzig Corpus Miner", a technical infrastructure for supporting qualitative and quantitative content analysis. The infrastructure aims at the integration of 'close reading' procedures on individual documents with…

Computation and Language · Computer Science 2017-07-12 Andreas Niekler , Gregor Wiedemann , Gerhard Heyer

This paper presents a comprehensive survey of corpora and lexical resources available for Turkish. We review a broad range of resources, focusing on the ones that are publicly available. In addition to providing information about the…

Computation and Language · Computer Science 2023-02-28 Çağrı Çöltekin , A. Seza Doğruöz , Özlem Çetinoğlu

Many computational argumentation tasks, like stance classification, are topic-dependent: the effectiveness of approaches to these tasks significantly depends on whether the approaches were trained on arguments from the same topics as those…

Computation and Language · Computer Science 2023-01-25 Yamen Ajjour , Johannes Kiesel , Benno Stein , Martin Potthast

Growing polarisation in society caught the attention of the scientific community as well as news media, which devote special issues to this phenomenon. At the same time, digitalisation of social interactions requires to revise concepts from…

Computation and Language · Computer Science 2024-08-13 Ewelina Gajewska , Katarzyna Budzynska , Barbara Konat , Marcin Koszowy , Konrad Kiljan , Maciej Uberna , He Zhang

A numeration system encodes abstract numeric quantities as concrete strings of written characters. The numeration systems used by modern scripts tend to be precise and unambiguous, but this was not so for the ancient and…

Computation and Language · Computer Science 2025-04-29 Logan Born , M. Willis Monroe , Kathryn Kelley , Anoop Sarkar

One of the most major and essential tasks in natural language processing is machine translation that is now highly dependent upon multilingual parallel corpora. Through this paper, we introduce the biggest Persian-English parallel corpus…

Computation and Language · Computer Science 2020-02-03 Omid Kashefi

A common practice in Natural Language Processing (NLP) is to visualize the text corpus without reading through the entire literature, still grasping the central idea and key points described. For a long time, researchers focused on…

Computation and Language · Computer Science 2022-07-29 Suvi Varshney , Divjeet Singh Jas

Visualization techniques have been widely used to analyze various data types, including text. This paper proposes an approach to analyze a controversial text in Portuguese by applying graph visualization techniques. Specifically, we use a…

Computers and Society · Computer Science 2024-06-24 Joao T. Aparicio , Andreas Karatsoli , Carlos J. Costa

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. \\…

Computation and Language · Computer Science 2026-04-29 B. van Es

The IMPACT-es diachronic corpus of historical Spanish compiles over one hundred books --containing approximately 8 million words-- in addition to a complementary lexicon which links more than 10 thousand lemmas with attestations of the…

Computation and Language · Computer Science 2013-07-01 Felipe Sánchez-Martínez , Isabel Martínez-Sempere , Xavier Ivars-Ribes , Rafael C. Carrasco

The process of debating is essential in our daily lives, whether in studying, work activities, simple everyday discussions, political debates on TV, or online discussions on social networks. The range of uses for debates is broad. Due to…