中文
相关论文

相关论文: Development of a Spanish Version of the Xerox Tagg…

200 篇论文

This working paper describes the Spanish tagset to be used in the context of CRATER, a CEC funded project aiming at the creation of a multilingual (English, French, Spanish) aligned corpus using the International Telecommunications Union…

cmp-lg · 计算机科学 2016-08-14 Fernando Sánchez León

This study explores the transfer learning capabilities of the TrOCR architecture to Spanish. TrOCR is a transformer-based Optical Character Recognition (OCR) model renowned for its state-of-the-art performance in English benchmarks.…

人工智能 · 计算机科学 2024-07-10 Filipe Lauar , Valentin Laurent

We present a bootstrapping method to develop an annotated corpus, which is specially useful for languages with few available resources. The method is being applied to develop a corpus of Spanish of over 5Mw. The method consists on taking…

计算与语言 · 计算机科学 2007-05-23 L. Marquez , L. Padro , H. Rodriguez

In this paper, we compare automated metrical pattern identification systems available for Spanish against extensive experiments done by fine-tuning language models trained on the same task. Despite being initially conceived as a model…

计算与语言 · 计算机科学 2020-11-20 Javier de la Rosa , Salvador Ros , Elena González-Blanco

This paper presents relevant issues that have been considered in the design of a general purpose lemmatizer/tagger for Basque (EUSLEM). The lemmatizer/tagger is conceived as a basic tool necessary for other linguistic applications. It uses…

The Spanish language is one of the top 5 spoken languages in the world. Nevertheless, finding resources to train or evaluate Spanish language models is not an easy task. In this paper we help bridge this gap by presenting a BERT-based…

计算与语言 · 计算机科学 2023-08-08 José Cañete , Gabriel Chaperon , Rodrigo Fuentes , Jou-Hui Ho , Hojin Kang , Jorge Pérez

The appearance of complex attention-based language models such as BERT, Roberta or GPT-3 has allowed to address highly complex tasks in a plethora of scenarios. However, when applied to specific domains, these models encounter considerable…

计算与语言 · 计算机科学 2022-06-14 Javier Huertas-Tato , Alejandro Martin , David Camacho

This paper presents a grammar and style checker demonstrator for Spanish and Greek native writers developed within the project GramCheck. Besides a brief grammar error typology for Spanish, a linguistically motivated approach to detection…

cmp-lg · 计算机科学 2016-08-15 Flora Ramírez Bustamante , Fernando Sánchez León

An overview of the present and foreseen R&D activities of the Spanish network for future accelerators aiming to participate in the design and construction of the forward tracker and vertex detectors of the Future Linear Colliders, is shown.

仪器与探测器 · 物理学 2010-06-16 Alberto Ruiz-Jimeno

This paper presents RigoBERTa, a State-of-the-Art Language Model for Spanish. RigoBERTa is trained over a well-curated corpus formed up from different subcorpora with key features. It follows the DeBERTa architecture, which has several…

Recognizing and categorizing lexical collocations in context is useful for language learning, dictionary compilation and downstream NLP. However, it is a challenging task due to the varying degrees of frozenness lexical collocations…

计算与语言 · 计算机科学 2022-05-24 Luis Espinosa-Anke , Alexander Shvets , Alireza Mohammadshahi , James Henderson , Leo Wanner

This work presents a new resource for borrowing identification and analyzes the performance and errors of several models on this task. We introduce a new annotated corpus of Spanish newswire rich in unassimilated lexical borrowings -- words…

计算与语言 · 计算机科学 2022-03-31 Elena Álvarez-Mellado , Constantine Lignos

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

We introduce CNER, an ensemble of capable tools for extraction of semantic relationships between named entities in Spanish language. Built upon a container-based architecture, CNER integrates different Named entity recognition and relation…

计算与语言 · 计算机科学 2024-05-20 Jefferson A. Peña Torres , Raúl E. Gutiérrez De Piñerez

A German language model for the Xerox HMM tagger is presented. This model's performance is compared with two other German taggers with partial parameter re-estimation and full adaption of parameters from pre-tagged corpora. The ambiguity…

cmp-lg · 计算机科学 2016-08-31 Helmut Feldweg

In this paper, we propose a methodology to align a medium-sized GPT model, originally trained in English for an open domain, to a small closed domain in Spanish. The application for which the model is finely tuned is the question answering…

计算与语言 · 计算机科学 2023-06-01 Oscar R. Navarrete-Parra , Victor Uc-Cetina , Jorge Reyes-Magana

In recent years, significant advancements in pre-trained language models have driven the creation of numerous non-English language variants, with a particular emphasis on encoder-only and decoder-only architectures. While Spanish language…

计算与语言 · 计算机科学 2024-03-22 Vladimir Araujo , Maria Mihaela Trusca , Rodrigo Tufiño , Marie-Francine Moens

We present the first shared task for detecting and analyzing code-switching in Guarani and Spanish, GUA-SPA at IberLEF 2023. The challenge consisted of three tasks: identifying the language of a token, NER, and a novel task of classifying…

This article presents the experiments and results obtained by the GRESEL team in the IberLEF 2025 shared task PastReader: Transcribing Texts from the Past. Three types of experiments were conducted with the dual aim of participating in the…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yanco Amor Torterolo-Orta , Jaione Macicior-Mitxelena , Marina Miguez-Lamanuzzi , Ana García-Serrano

The relationship between language model tokenization and performance is an open area of research. Here, we investigate how different tokenization schemes impact number agreement in Spanish plurals. We find that morphologically-aligned…

计算与语言 · 计算机科学 2024-03-21 Catherine Arnett , Pamela D. Rivière , Tyler A. Chang , Sean Trott
‹ 上一页 1 2 3 10 下一页 ›