English
Related papers

Related papers: A Spanish Tagset for the CRATER Project

200 papers

Document retrieval has greatly benefited from the advancements of large-scale pre-trained language models (PLMs). However, their effectiveness is often limited in theme-specific applications for specialized areas or industries, due to…

Information Retrieval · Computer Science 2024-03-08 SeongKu Kang , Shivam Agarwal , Bowen Jin , Dongha Lee , Hwanjo Yu , Jiawei Han

PaECTER is an open-source document-level encoder specific for patents. We fine-tune BERT for Patents with examiner-added citation information to generate numerical representations for patent documents. PaECTER performs better in similarity…

Information Retrieval · Computer Science 2025-10-02 Mainak Ghosh , Michael E. Rose , Sebastian Erhardt , Erik Buunk , Dietmar Harhoff

Legal corpora for Natural Language Processing (NLP) are valuable and scarce resources in languages like Spanish due to two main reasons: data accessibility and legal expert knowledge availability. INESData 2024 is a European Union funded…

We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century. In a series of case studies, we illustrate the affordances…

Computation and Language · Computer Science 2020-09-22 David Bamman , Patrick J. Burns

We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages. We then compare the…

Computation and Language · Computer Science 2020-08-24 Pedro Javier Ortiz Suárez , Laurent Romary , Benoît Sagot

Cross-lingual annotations of legislative texts enable us to explore major themes covered in multilingual legal data and are a key facilitator of semantic similarity when searching for similar documents. Multilingual probabilistic topic…

Information Retrieval · Computer Science 2019-12-02 Carlos Badenes-Olmedo , Jose-Luis Redondo-Garcia , Oscar Corcho

The popularity of social media has created problems such as hate speech and sexism. The identification and classification of sexism in social media are very relevant tasks, as they would allow building a healthier social environment.…

Computation and Language · Computer Science 2021-11-09 Angel Felipe Magnossão de Paula , Roberto Fray da Silva , Ipek Baris Schlicht

Homonym identification is important for WSD that require coarse-grained partitions of senses. The goal of this project is to determine whether contextual information is sufficient for identifying a homonymous word. To capture the context,…

Computation and Language · Computer Science 2021-01-08 Rohan Saha

This paper presents new state-of-the-art models for three tasks, part-of-speech tagging, syntactic parsing, and semantic parsing, using the cutting-edge contextualized embedding framework known as BERT. For each task, we first replicate and…

Computation and Language · Computer Science 2020-05-26 Han He , Jinho D. Choi

We present a submission to the CogALex 2016 shared task on the corpus-based identification of semantic relations, using LexNET (Shwartz and Dagan, 2016), an integrated path-based and distributional method for semantic relation…

Computation and Language · Computer Science 2016-11-02 Vered Shwartz , Ido Dagan

Learning fine-grained distinctions between vocabulary items is a key challenge in learning a new language. For example, the noun "wall" has different lexical manifestations in Spanish -- "pared" refers to an indoor wall while "muro" refers…

Computation and Language · Computer Science 2021-09-14 Aditi Chaudhary , Kayo Yin , Antonios Anastasopoulos , Graham Neubig

Lexicon-based methods using syntactic rules for polarity classification rely on parsers that are dependent on the language and on treebank guidelines. Thus, rules are also dependent and require adaptation, especially in multilingual…

Computation and Language · Computer Science 2017-08-18 David Vilares , Marcos Garcia , Miguel A. Alonso , Carlos Gómez-Rodríguez

In this paper we will present our ongoing work on a plan-based discourse processor developed in the context of the Enthusiast Spanish to English translation system as part of the JANUS multi-lingual speech-to-speech translation system. We…

cmp-lg · Computer Science 2008-02-03 Carolyn Penstein Rose' , Barbara Di Eugenio , Lori S. Levin , Carol Van Ess-Dykema

Text simplification plays a crucial role in improving the accessibility and comprehensibility of written information for diverse audiences, including language learners and readers with limited literacy. Despite its importance, large-scale,…

Computation and Language · Computer Science 2026-05-12 Kenji Hilasaca , Nouran Khallaf , Serge Sharoff

Person-job fit is an essential part of online recruitment platforms in serving various downstream applications like Job Search and Candidate Recommendation. Recently, pretrained large language models have further enhanced the effectiveness…

Computation and Language · Computer Science 2024-01-19 Yihan Cao , Xu Chen , Lun Du , Hao Chen , Qiang Fu , Shi Han , Yushu Du , Yanbin Kang , Guangming Lu , Zi Li

We present the first shared task on semantic change discovery and detection in Spanish and create the first dataset of Spanish words manually annotated for semantic change using the DURel framework (Schlechtweg et al., 2018). The task is…

Computation and Language · Computer Science 2022-05-16 Frank D. Zamora-Reina , Felipe Bravo-Marquez , Dominik Schlechtweg

We present a new corpus of Twitter data annotated for codeswitching and borrowing between Spanish and English. The corpus contains 9,500 tweets annotated at the token level with codeswitches, borrowings, and named entities. This corpus…

Computation and Language · Computer Science 2022-06-13 Elena Alvarez Mellado , Constantine Lignos

This paper analyzes multiple deep-syntactic frameworks with the goal of creating a proposal for a set of universal semantic role labels. The proposal examines various theoretic linguistic perspectives and focuses on Meaning-Text Theory and…

Computation and Language · Computer Science 2023-03-23 Kira Droganova , Daniel Zeman

In this paper we consider two sequence tagging tasks for medieval Latin: part-of-speech tagging and lemmatization. These are both basic, yet foundational preprocessing steps in applications such as text re-use detection. Nevertheless, they…

Computation and Language · Computer Science 2023-06-22 Mike Kestemont , Jeroen De Gussem

Existing Latin treebanks draw from Latin's long written tradition, spanning 17 centuries and a variety of cultures. Recent efforts have begun to harmonize these treebanks' annotations to better train and evaluate morphological taggers.…

Computation and Language · Computer Science 2024-08-14 Marisa Hudspeth , Brendan O'Connor , Laure Thompson
‹ Prev 1 3 4 5 6 7 10 Next ›