English

Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers

Computation and Language 2024-05-14 v1

Abstract

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimodal data, which includes texts, tables and figures. The paper presents the results of experiments with two language models: Gigachat from SBER and YandexGPT from Yandex. The dataset consists of 420 papers and is publicly available on https://github.com/iis-research-team/summarization-dataset.

Keywords

Cite

@article{arxiv.2405.07886,
  title  = {Russian-Language Multimodal Dataset for Automatic Summarization of Scientific Papers},
  author = {Alena Tsanda and Elena Bruches},
  journal= {arXiv preprint arXiv:2405.07886},
  year   = {2024}
}

Comments

12 pages, accepted to AINL

R2 v1 2026-06-28T16:25:36.734Z