English
Related papers

Related papers: An Annotated Corpus of Emerging Anglicisms in Span…

200 papers

This paper summarizes the main findings of ADoBo 2025, the shared task on anglicism identification in Spanish proposed in the context of IberLEF 2025. Participants of ADoBo 2025 were asked to detect English lexical borrowings (or…

Computation and Language · Computer Science 2025-07-30 Elena Alvarez-Mellado , Jordi Porta-Zamorano , Constantine Lignos , Julio Gonzalo

This work presents a new resource for borrowing identification and analyzes the performance and errors of several models on this task. We introduce a new annotated corpus of Spanish newswire rich in unassimilated lexical borrowings -- words…

Computation and Language · Computer Science 2022-03-31 Elena Álvarez-Mellado , Constantine Lignos

This paper summarizes the main findings of the ADoBo 2021 shared task, proposed in the context of IberLef 2021. In this task, we invited participants to detect lexical borrowings (coming mostly from English) in Spanish newswire texts. This…

Computation and Language · Computer Science 2021-11-01 Elena Álvarez Mellado , Luis Espinosa Anke , Julio Gonzalo Arroyo , Constantine Lignos , Jordi Porta Zamorano

We develop novel annotation guidelines for sentence-level subjectivity detection, which are not limited to language-specific cues. We use our guidelines to collect NewsSD-ENG, a corpus of 638 objective and 411 subjective sentences extracted…

Newspapers are a popular form of written discourse, read by many people, thanks to the novelty of the information provided by the news content in it. A headline is the most widely read part of any newspaper due to its appearance in a bigger…

Computation and Language · Computer Science 2019-10-21 Elizabeth Jasmi George , Radhika Mamidi

The Abstract Meaning Representation (AMR) formalism, designed originally for English, has been adapted to a number of languages. We build on previous work proposing the annotation of AMR in Spanish, which resulted in the release of 50…

Computation and Language · Computer Science 2022-04-19 Shira Wein , Lucia Donatelli , Ethan Ricker , Calvin Engstrom , Alex Nelson , Nathan Schneider

English news headlines form a register with unique syntactic properties that have been documented in linguistics literature since the 1930s. However, headlines have received surprisingly little attention from the NLP syntactic parsing…

Computation and Language · Computer Science 2023-01-26 Adrian Benton , Tianze Shi , Ozan İrsoy , Igor Malioutov

This paper presents a new annotated corpus of 513 anonymized radiology reports written in Spanish. Reports were manually annotated with entities, negation and uncertainty terms and relations. The corpus was conceived as an evaluation…

Computation and Language · Computer Science 2017-11-01 Viviana Cotik , Darío Filippo , Roland Roller , Hans Uszkoreit , Feiyu Xu

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

Computation and Language · Computer Science 2023-08-25 Emily Silcock , Melissa Dell

The IMPACT-es diachronic corpus of historical Spanish compiles over one hundred books --containing approximately 8 million words-- in addition to a complementary lexicon which links more than 10 thousand lemmas with attestations of the…

Computation and Language · Computer Science 2013-07-01 Felipe Sánchez-Martínez , Isabel Martínez-Sempere , Xavier Ivars-Ribes , Rafael C. Carrasco

NLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible. Developing such methods poses substantial challenges though. First, acquiring large, annotated historical datasets is difficult, as…

Computation and Language · Computer Science 2023-05-19 Nadav Borenstein , Natalia da Silva Perez , Isabelle Augenstein

Event extraction is an NLP task that commonly involves identifying the central word (trigger) for an event and its associated arguments in text. ACE-2005 is widely recognised as the standard corpus in this field. While other corpora, like…

Computation and Language · Computer Science 2024-09-02 Luís Filipe Cunha , Purificação Silvano , Ricardo Campos , Alípio Jorge

Quotation extraction is a widely useful task both from a sociological and from a Natural Language Processing perspective. However, very little data is available to study this task in languages other than English. In this paper, we present a…

Computation and Language · Computer Science 2023-09-20 Ange Richard , Laura Alonzo-Canul , François Portet

Acronym extraction is the task of identifying acronyms and their expanded forms in texts that is necessary for various NLP applications. Despite major progress for this task in recent years, one limitation of existing AE research is that…

Computation and Language · Computer Science 2022-02-22 Amir Pouran Ben Veyseh , Nicole Meister , Seunghyun Yoon , Rajiv Jain , Franck Dernoncourt , Thien Huu Nguyen

In this paper, we present InferES - an original corpus for Natural Language Inference (NLI) in European Spanish. We propose, implement, and analyze a variety of corpus-creating strategies utilizing expert linguists and crowd workers. The…

Computation and Language · Computer Science 2022-10-07 Venelin Kovatchev , Mariona Taulé

We describe a rule-based approach for the automatic acquisition of salient scientific entities from Computational Linguistics (CL) scholarly article titles. Two observations motivated the approach: (i) noting salient aspects of an article's…

Information Retrieval · Computer Science 2021-09-20 Jennifer D'Souza , Soeren Auer

This work introduces a machine translation task where the output is aimed at audiences of different levels of target language proficiency. We collect a high quality dataset of news articles available in English and Spanish, written for…

Computation and Language · Computer Science 2019-11-05 Sweta Agrawal , Marine Carpuat

This paper presents two significant contributions: First, it introduces a novel dataset of 19th-century Latin American newspaper texts, addressing a critical gap in specialized corpora for historical and linguistic analysis in this region.…

Computation and Language · Computer Science 2025-03-31 Laura Manrique-Gómez , Tony Montes , Arturo Rodríguez-Herrera , Rubén Manrique

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are very few Spanish Language Models which result to be small and…

Computation and Language · Computer Science 2021-10-26 Asier Gutiérrez-Fandiño , Jordi Armengol-Estapé , Aitor Gonzalez-Agirre , Marta Villegas

The popularity of automated news headline generation has surged with advancements in pre-trained language models. However, these models often suffer from the ``hallucination'' problem, where the generated headline is not fully supported by…

Computation and Language · Computer Science 2024-07-24 Jiaming Shen , Tianqi Liu , Jialu Liu , Zhen Qin , Jay Pavagadhi , Simon Baumgartner , Michael Bendersky
‹ Prev 1 2 3 10 Next ›