English
Related papers

Related papers: AGRR-2019: A Corpus for Gapping Resolution in Russ…

200 papers

The quality and accessibility of multilingual datasets are crucial for advancing machine translation. However, previous corpora built from United Nations documents have suffered from issues such as opaque process, difficulty of…

Computation and Language · Computer Science 2025-09-22 Qiuyang Lu , Fangjian Shen , Zhengkai Tang , Qiang Liu , Hexuan Cheng , Hui Liu , Wushao Wen

ChatGPT is a conversational artificial intelligence that is a member of the generative pre-trained transformer of the large language model family. This text generative model was fine-tuned by both supervised learning and reinforcement…

Computation and Language · Computer Science 2023-06-06 Niful Islam , Debopom Sutradhar , Humaira Noor , Jarin Tasnim Raya , Monowara Tabassum Maisha , Dewan Md Farid

In this paper, we present a new Russian and Kazakh database (with about 95% of Russian and 5% of Kazakh words/sentences respectively) for offline handwriting recognition. A few pre-processing and segmentation procedures have been developed…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Daniyar Nurseitov , Kairat Bostanbekov , Daniyar Kurmankhojayev , Anel Alimova , Abdelrahman Abdallah

With the growing interest in large language models, the need for evaluating the quality of machine text compared to reference (typically human-generated) text has become focal attention. Most recent works focus either on task-specific…

Automatic text generation based on neural language models has achieved performance levels that make the generated text almost indistinguishable from those written by humans. Despite the value that text generation can have in various…

Computation and Language · Computer Science 2022-05-02 Vijini Liyanage , Davide Buscaldi , Adeline Nazarenko

Generative large language models (LLMs) have become crucial for modern NLP research and applications across various languages. However, the development of foundational models specifically tailored to the Russian language has been limited,…

In this paper, we present a distributional word embedding model trained on one of the largest available Russian corpora: Araneum Russicum Maximum (over 10 billion words crawled from the web). We compare this model to the model trained on…

Computation and Language · Computer Science 2018-01-22 Andrey Kutuzov , Maria Kunilovskaya

Automatic text summarization has been studied in a variety of domains and languages. However, this does not hold for the Russian language. To overcome this issue, we present Gazeta, the first dataset for summarization of Russian news. We…

Computation and Language · Computer Science 2021-10-06 Ilya Gusev

We present a new data set for speech emotion recognition (SER) tasks called Dusha. The corpus contains approximately 350 hours of data, more than 300 000 audio recordings with Russian speech and their transcripts. Therefore it is the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-26 Vladimir Kondratenko , Artem Sokolov , Nikolay Karpov , Oleg Kutuzov , Nikita Savushkin , Fyodor Minkin

The Split and Rephrase (SPRP) task, which consists in splitting complex sentences into a sequence of shorter grammatical sentences, while preserving the original meaning, can facilitate the processing of complex texts for humans and…

Computation and Language · Computer Science 2024-10-11 David Ponce , Thierry Etchegoyhen , Jesús Calleja Pérez , Harritxu Gete

The main goal of this paper is to develop a spell checker module for clinical text in Russian. The described approach combines string distance measure algorithms with technics of machine learning embedding methods. Our overall precision is…

Computation and Language · Computer Science 2020-07-30 Ksenia Balabaeva , Anastasia Funkner , Sergey Kovalchuk

In the last year, new neural architectures and multilingual pre-trained models have been released for Russian, which led to performance evaluation problems across a range of language understanding tasks. This paper presents Russian…

Grammatical error correction is one of the fundamental tasks in Natural Language Processing. For the Russian language, most of the spellcheckers available correct typos and other simple errors with high accuracy, but often fail when faced…

Computation and Language · Computer Science 2023-07-06 Nikita Remnev , Sergei Obiedkov , Ekaterina Rakhilina , Ivan Smirnov , Anastasia Vyrenkova

We present 20min-XD (20 Minuten cross-lingual document-level), a French-German, document-level comparable corpus of news articles, sourced from the Swiss online news outlet 20 Minuten/20 minutes. Our dataset comprises around 15,000 article…

Computation and Language · Computer Science 2025-05-01 Michelle Wastl , Jannis Vamvas , Selena Calleri , Rico Sennrich

The paper introduces manually annotated test sets for the task of tracing diachronic (temporal) semantic shifts in Russian. The two test sets are complementary in that the first one covers comparatively strong semantic changes occurring to…

Computation and Language · Computer Science 2019-07-31 Vadim Fomin , Daria Bakshandaeva , Julia Rodina , Andrey Kutuzov

This paper presents an overview of rule-based system for automatic accentuation and phonemic transcription of Russian texts for speech connected tasks, such as Automatic Speech Recognition (ASR). Two parts of the developed system,…

Computation and Language · Computer Science 2024-10-07 Olga Iakovenko , Ivan Bondarenko , Mariya Borovikova , Daniil Vodolazsky

In this paper, we provide an account of how to generate sentences with coordination constructions from clause-sized semantic representations. An algorithm is developed to generate sentences with ellipsis, gapping, right-node-raising, and…

cmp-lg · Computer Science 2007-05-23 James Shaw

To explain social phenomena and identify systematic biases, much research in computational social science focuses on comparative text analyses. These studies often rely on coarse corpus-level statistics or local word-level analyses, mainly…

Computation and Language · Computer Science 2024-10-08 Farhan Samir , Chan Young Park , Anjalie Field , Vered Shwartz , Yulia Tsvetkov

Understanding stories is a challenging reading comprehension problem for machines as it requires reading a large volume of text and following long-range dependencies. In this paper, we introduce the Shmoop Corpus: a dataset of 231 stories…

Computation and Language · Computer Science 2020-01-03 Atef Chaudhury , Makarand Tapaswi , Seung Wook Kim , Sanja Fidler

We present DNIPRO, a corpus of 246K news articles from the Russo-Ukrainian war (Feb 2022 -- Aug 2024) spanning eleven outlets across five nation-states (Russia, Ukraine, U.S., U.K., China) and three languages. The corpus features…

Computation and Language · Computer Science 2026-03-16 Dikshya Mohanty , Taisiia Sabadyn , Jelwin Rodrigues , Chenlu Wang , Abhishek Kalugade , Ritwik Banerjee