中文
相关论文

相关论文: RuDSI: graph-based word sense induction dataset fo…

200 篇论文

In this paper, we present Watasense, an unsupervised system for word sense disambiguation. Given a sentence, the system chooses the most relevant sense of each input word with respect to the semantic similarity between the given sentence…

Generating coherent, grammatically correct, and meaningful text is very challenging, however, it is crucial to many modern NLP systems. So far, research has mostly focused on English language, for other languages both standardized datasets,…

计算与语言 · 计算机科学 2020-05-07 Zein Shaheen , Gerhard Wohlgenannt , Bassel Zaity , Dmitry Mouromtsev , Vadim Pak

This paper presents an overview of rule-based system for automatic accentuation and phonemic transcription of Russian texts for speech connected tasks, such as Automatic Speech Recognition (ASR). Two parts of the developed system,…

计算与语言 · 计算机科学 2024-10-07 Olga Iakovenko , Ivan Bondarenko , Mariya Borovikova , Daniil Vodolazsky

Radiology reports play a critical role in communicating medical findings to physicians. In each report, the impression section summarizes essential radiology findings. In clinical practice, writing impression is highly demanded yet…

计算与语言 · 计算机科学 2021-12-21 Jinpeng Hu , Jianling Li , Zhihong Chen , Yaling Shen , Yan Song , Xiang Wan , Tsung-Hui Chang

Logical rules are a popular knowledge representation language in many domains, representing background knowledge and encoding information that can be derived from given facts in a compact form. However, rule formulation is a complex process…

人工智能 · 计算机科学 2020-02-13 Cristina Cornelio , Veronika Thost

In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to text-to-speech…

计算与语言 · 计算机科学 2024-10-01 Tomilov A. A. , Gromova A. Y. , Svischev A. N

In this paper, we present a novel series of Russian information retrieval datasets constructed from the "Did you know..." section of Russian Wikipedia. Our datasets support a range of retrieval tasks, including fact-checking,…

信息检索 · 计算机科学 2025-11-10 Grigory Kovalev , Natalia Loukachevitch , Mikhail Tikhomirov , Olga Babina , Pavel Mamaev

In this paper, we introduce an advanced Russian general language understanding evaluation benchmark -- RussianGLUE. Recent advances in the field of universal language models and transformers require the development of a methodology for…

The quality of natural language texts in fine-tuning datasets plays a critical role in the performance of generative models, particularly in computational creativity tasks such as poem or song lyric generation. Fluency defects in generated…

计算与语言 · 计算机科学 2025-05-08 Ilya Koziev

In the last year, new neural architectures and multilingual pre-trained models have been released for Russian, which led to performance evaluation problems across a range of language understanding tasks. This paper presents Russian…

Currently, there are more than a dozen Russian-language corpora for sentiment analysis, differing in the source of the texts, domain, size, number and ratio of sentiment classes, and annotation method. This work examines publicly available…

计算与语言 · 计算机科学 2021-06-29 Evgeny Kotelnikov

In this paper, we present Russian language datasets in the digital humanities domain for the evaluation of word embedding techniques or similar language modeling and feature learning algorithms. The datasets are split into two task types,…

计算与语言 · 计算机科学 2019-03-22 Gerhard Wohlgenannt , Artemii Babushkin , Denis Romashov , Igor Ukrainets , Anton Maskaykin , Ilya Shutov

Generative poetry systems require effective tools for data engineering and automatic evaluation, particularly to assess how well a poem adheres to versification rules, such as the correct alternation of stressed and unstressed syllables and…

计算与语言 · 计算机科学 2025-10-21 Ilya Koziev

WordNet-like Lexical Databases (WLDs) group English words into sets of synonyms called "synsets." Although the standard WLDs are being used in many successful Text-Mining applications, they have the limitation that word-senses are…

This study investigates the feasibility of automating clinical coding in Russian, a language with limited biomedical resources. We present a new dataset for ICD coding, which includes diagnosis fields from electronic health records (EHRs)…

Information surrounds people in modern life. Text is a very efficient type of information that people use for communication for centuries. However, automated text-in-the-wild recognition remains a challenging problem. The major limitation…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Igor Markov , Sergey Nesteruk , Andrey Kuznetsov , Denis Dimitrov

Word Sense Disambiguation (WSD), the process of automatically identifying the meaning of a polysemous word in a sentence, is a fundamental task in Natural Language Processing (NLP). Progress in this approach to WSD opens up many promising…

计算与语言 · 计算机科学 2013-10-08 Mohammad Nasiruddin

Word sense induction (WSI), or the task of automatically discovering multiple senses or meanings of a word, has three main challenges: domain adaptability, novel sense detection, and sense granularity flexibility. While current latent…

计算与语言 · 计算机科学 2018-11-26 Reinald Kim Amplayo , Seung-won Hwang , Min Song

The goal of this work is to improve the performance of a neural named entity recognition system by adding input features that indicate a word is part of a name included in a gazetteer. This article describes how to generate gazetteers from…

计算与语言 · 计算机科学 2020-03-09 Chan Hee Song , Dawn Lawrie , Tim Finin , James Mayfield

Conventional word sense induction (WSI) methods usually represent each instance with discrete linguistic features or cooccurrence features, and train a model for each polysemous word individually. In this work, we propose to learn sense…

计算与语言 · 计算机科学 2016-06-23 Linfeng Song , Zhiguo Wang , Haitao Mi , Daniel Gildea