English
Related papers

Related papers: A Spanish Tagset for the CRATER Project

200 papers

The paper proposes the task of universal semantic tagging---tagging word tokens with language-neutral, semantically informative tags. We argue that the task, with its independent nature, contributes to better semantic analysis for…

Computation and Language · Computer Science 2017-10-02 Lasha Abzianidze , Johan Bos

An overview of the present and foreseen R&D activities of the Spanish network for future accelerators aiming to participate in the design and construction of the forward tracker and vertex detectors of the Future Linear Colliders, is shown.

Instrumentation and Detectors · Physics 2010-06-16 Alberto Ruiz-Jimeno

Code-switching, or alternating between languages within a single conversation, presents challenges for multilingual language models on NLP tasks. This research investigates if pre-training Multilingual BERT (mBERT) on code-switched datasets…

Computation and Language · Computer Science 2025-03-12 Katherine Xie , Nitya Babbar , Vicky Chen , Yoanna Turura

We describe the design and use of the CREER dataset, a large corpus annotated with rich English grammar and semantic attributes. The CREER dataset uses the Stanford CoreNLP Annotator to capture rich language structures from Wikipedia plain…

Computation and Language · Computer Science 2022-06-10 Yu-Siou Tang , Chung-Hsien Wu

Diacritic characters can be considered as a unique set of characters providing us with adequate and significant clue in identifying a given language with considerably high accuracy. Diacritics, though associated with phonetics often serve…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Shubham Vatsal , Nikhil Arora , Gopi Ramena , Sukumar Moharana , Dhruval Jain , Naresh Purre , Rachit S Munjal

This paper maps the large-scale variation of the Spanish language by employing a corpus based on geographically tagged Twitter messages. Lexical dialects are extracted from an analysis of variants of tens of concepts. The resulting maps…

Machine Learning · Statistics 2017-02-07 Bruno Gonçalves , David Sánchez

We describe a set of bilingual English--French and English--German parallel corpora in which the direction of translation is accurately and reliably annotated. The corpora are diverse, consisting of parliamentary proceedings, literary…

Computation and Language · Computer Science 2016-03-08 Ella Rabinovich , Shuly Wintner , Ofek Luis Lewinsohn

We present an extended comparison of contextualized language models for Hungarian. We compare huBERT, a Hungarian model against 4 multilingual models including the multilingual BERT model. We evaluate these models through three tasks,…

Computation and Language · Computer Science 2021-02-23 Judit Ács , Dániel Lévai , Dávid Márk Nemeskey , András Kornai

We introduce a novel dependency parser, the hexatagger, that constructs dependency trees by tagging the words in a sentence with elements from a finite set of possible tags. In contrast to many approaches to dependency parsing, our approach…

Computation and Language · Computer Science 2023-08-01 Afra Amini , Tianyu Liu , Ryan Cotterell

The goal of our project is to develop an accurate tagger for questions posted on Stack Exchange. Our problem is an instance of the more general problem of developing accurate classifiers for large scale text datasets. We are tackling the…

Computation and Language · Computer Science 2015-12-15 Sanket Mehta , Shagun Sodhani

This paper introduces LatinCy, a set of trained general purpose Latin-language "core" pipelines for use with the spaCy natural language processing framework. The models are trained on a large amount of available Latin data, including all…

Computation and Language · Computer Science 2023-05-09 Patrick J. Burns

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are very few Spanish Language Models which result to be small and…

Computation and Language · Computer Science 2021-10-26 Asier Gutiérrez-Fandiño , Jordi Armengol-Estapé , Aitor Gonzalez-Agirre , Marta Villegas

This paper describes Charles University submission for Terminology translation Shared Task at WMT21. The objective of this task is to design a system which translates certain terms based on a provided terminology database, while preserving…

Computation and Language · Computer Science 2021-09-21 Josef Jon , Michal Novák , João Paulo Aires , Dušan Variš , Ondřej Bojar

Multilingual language models have been a crucial breakthrough as they considerably reduce the need of data for under-resourced languages. Nevertheless, the superiority of language-specific models has already been proven for languages having…

In this paper we introduce the methodology used and the basic phases we followed to develop the Catalan WordNet, and shich lexical resources have been employed in its building. This methodology, as well as the tools we made use of, have…

cmp-lg · Computer Science 2007-05-23 Laura Benitez , Sergi Cervell , Gerard Escudero , Monica Lopez , German Rigau , Mariona Taule

This paper introduces an LLM-based Latin-to-English translation platform designed to address the challenges of translating Latin texts. We named the model LITERA, which stands for Latin Interpretation and Translations into English for…

Computation and Language · Computer Science 2025-04-16 Paul Rosu

Named Entity Recognition (NER) is a critical component of Natural Language Processing (NLP) for extracting structured information from unstructured text. However, for low-resource languages like Catalan, the performance of NER systems often…

The paper presents the Source Code Analysis and Lexical Annotation Runtime (SCALAR), a tool specialized for mapping (annotating) source code identifier names to their corresponding part-of-speech tag sequence (grammar pattern). SCALAR's…

In this report, we present TAGLAS, an atlas of text-attributed graph (TAG) datasets and benchmarks. TAGs are graphs with node and edge features represented in text, which have recently gained wide applicability in training graph-language or…

Machine Learning · Computer Science 2024-10-22 Jiarui Feng , Hao Liu , Lecheng Kong , Mingfang Zhu , Yixin Chen , Muhan Zhang

We address the problem of Part of Speech tagging (POS) in the context of linguistic code switching (CS). CS is the phenomenon where a speaker switches between two languages or variants of the same language within or across utterances, known…

Computation and Language · Computer Science 2019-11-05 Fahad AlGhamdi , Giovanni Molina , Mona Diab , Thamar Solorio , Abdelati Hawwari , Victor Soto , Julia Hirschberg