中文
相关论文

相关论文: A Publicly Available Cross-Platform Lemmatizer for…

200 篇论文

This paper reveals the results of an analysis of the accuracy of developed software for automatic lemmatization for the Bulgarian language. This lemmatization software is written entirely in Java and is distributed as a GATE plugin. Certain…

计算与语言 · 计算机科学 2015-06-16 Elena Karashtranova , Grigor Iliev , Nadezhda Borisova , Yana Chankova , Irena Atanasova

In this article, we describe an approach for automatic detection of noun-adjective agreement errors in Bulgarian texts by explaining the necessary steps required to develop a simple Java-based language processing application. For this…

计算与语言 · 计算机科学 2014-11-04 Nadezhda Borisova , Grigor Iliev , Elena Karashtranova

We present experiments with part-of-speech tagging for Bulgarian, a Slavic language with rich inflectional and derivational morphology. Unlike most previous work, which has used a small number of grammatical categories, we work with 680…

计算与语言 · 计算机科学 2019-11-27 Georgi Georgiev , Valentin Zhikov , Petya Osenova , Kiril Simov , Preslav Nakov

We present BgGPT-Gemma-2-27B-Instruct and BgGPT-Gemma-2-9B-Instruct: continually pretrained and fine-tuned versions of Google's Gemma-2 models, specifically optimized for Bulgarian language understanding and generation. Leveraging Gemma-2's…

计算与语言 · 计算机科学 2024-12-17 Anton Alexandrov , Veselin Raychev , Dimitar I. Dimitrov , Ce Zhang , Martin Vechev , Kristina Toutanova

In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently a lack of available…

计算与语言 · 计算机科学 2019-12-03 Nelda Kote , Marenglen Biba , Jenna Kanerva , Samuel Rönnqvist , Filip Ginter

In this paper we present Morphy, an integrated tool for German morphology, part-of-speech tagging and context-sensitive lemmatization. Its large lexicon of more than 320,000 word forms plus its ability to process German compound nouns…

计算与语言 · 计算机科学 2007-05-23 Wolfgang Lezius , Reinhard Rapp , Manfred Wettler

We present bgGLUE(Bulgarian General Language Understanding Evaluation), a benchmark for evaluating language models on Natural Language Understanding (NLU) tasks in Bulgarian. Our benchmark includes NLU tasks targeting a variety of NLP…

We present GliLem -- a novel hybrid lemmatization system for Estonian that enhances the highly accurate rule-based morphological analyzer Vabamorf with an external disambiguation module based on GliNER -- an open vocabulary NER model that…

计算与语言 · 计算机科学 2025-01-14 Aleksei Dorkin , Kairit Sirts

Toxic content detection in online communication remains a significant challenge, with current solutions often inadvertently blocking valuable information, including medical terms and text related to minority groups. This paper presents a…

计算与语言 · 计算机科学 2026-04-03 Melania Berbatova , Tsvetoslav Vasev

This paper presents a corpus manually annotated with named entities for six Slavic languages - Bulgarian, Czech, Polish, Slovenian, Russian, and Ukrainian. This work is the result of a series of shared tasks, conducted in 2017-2023 as a…

计算与语言 · 计算机科学 2024-04-09 Jakub Piskorski , Michał Marcińczuk , Roman Yangarber

Lemmatization -- the task of mapping an inflected word form to its dictionary form -- is a crucial component of many NLP applications. In this paper, we present RUMLEM, a lemmatizer that covers the five main varieties of Romansh as well as…

计算与语言 · 计算机科学 2026-04-14 Dominic P. Fischer , Zachary Hopton , Jannis Vamvas

The paper presents a feature-rich approach to the automatic recognition and categorization of named entities (persons, organizations, locations, and miscellaneous) in news text for Bulgarian. We combine well-established features used for…

计算与语言 · 计算机科学 2021-10-01 Georgi Georgiev , Preslav Nakov , Kuzman Ganchev , Petya Osenova , Kiril Ivanov Simov

Lemmatization, finding the basic morphological form of a word in a corpus, is an important step in many natural language processing tasks when working with morphologically rich languages. We describe and evaluate Nefnir, a new open source…

Lemmatization holds significance in both natural language processing (NLP) and linguistics, as it effectively decreases data density and aids in comprehending contextual meaning. However, due to the highly inflected nature and morphological…

The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for…

计算与语言 · 计算机科学 2025-10-13 Stefan Krsteski , Matea Tashkovska , Borjan Sazdov , Hristijan Gjoreski , Branislav Gerazov

Large language models (LLMs) have become an essential tool for natural language processing and artificial intelligence in general. Current open-source models are primarily trained on English texts, resulting in poorer performance on…

计算与语言 · 计算机科学 2026-03-03 Domen Vreš , Tjaša Arčon , Timotej Petrič , Dario Vajda , Marko Robnik-Šikonja , Iztok Lebar Bajec

We present LEMMING, a modular log-linear model that jointly models lemmatization and tagging and supports the integration of arbitrary global features. It is trainable on corpora annotated with gold standard tags and lemmata and does not…

计算与语言 · 计算机科学 2024-05-29 Thomas Muller , Ryan Cotterell , Alexander Fraser , Hinrich Schütze

Lemmatization is the process of grouping together the inflected forms of a word so they can be analysed as a single item, identified by the word's lemma, or dictionary form. In computational linguistics, lemmatisation is the algorithmic…

计算与语言 · 计算机科学 2022-07-26 Michal Karwatowski , Marcin Pietron

Lemmatization is one of the core concepts in natural language processing, thus creating a lemmatization tool is an important task. This paper discusses the construction of a lemmatization algorithm for the Uzbek language. The main purpose…

计算与语言 · 计算机科学 2022-10-31 Maksud Sharipov , Ogabek Sobirov

Recently, reading comprehension models achieved near-human performance on large-scale datasets such as SQuAD, CoQA, MS Macro, RACE, etc. This is largely due to the release of pre-trained contextualized representations such as BERT and ELMo,…

计算与语言 · 计算机科学 2019-09-09 Momchil Hardalov , Ivan Koychev , Preslav Nakov
‹ 上一页 1 2 3 10 下一页 ›