English
Related papers

Related papers: Analysis of the quotation corpus of the Russian Wi…

200 papers

We present RUSLAN -- a new open Russian spoken language corpus for the text-to-speech task. RUSLAN contains 22200 audio samples with text annotations -- more than 31 hours of high-quality speech of one person -- being the largest annotated…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-28 Lenar Gabdrakhmanov , Rustem Garaev , Evgenii Razinkov

Semantic relatedness of terms represents similarity of meaning by a numerical score. On the one hand, humans easily make judgments about semantic relatedness. On the other hand, this kind of information is useful in language processing…

This paper provides a comprehensive overview of the gapping dataset for Russian that consists of 7.5k sentences with gapping (as well as 15k relevant negative sentences) and comprises data from various genres: news, fiction, social media…

Computation and Language · Computer Science 2019-06-11 Maria Ponomareva , Kira Droganova , Ivan Smurov , Tatiana Shavrina

A major challenge in paraphrase research is the lack of parallel corpora. In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs. The main advantage of our…

Computation and Language · Computer Science 2017-08-02 Wuwei Lan , Siyu Qiu , Hua He , Wei Xu

We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905 tokens). The corpus is available in two versions: the basic version contains texts with…

Computation and Language · Computer Science 2026-04-29 Denis Saveliev , Ruslan Kuchakov

We present a set of deterministic algorithms for Russian inflection and automated text synthesis. These algorithms are implemented in a publicly available web-service www.passare.ru. This service provides functions for inflection of single…

Computation and Language · Computer Science 2023-06-02 A. A. Gurin , T. M. Sadykov , T. A. Zhukov

Corpora that contain tabular data such as WebTables are a vital resource for the academic community. Essentially, they are the backbone of any modern research in information management. They are used for various tasks of data extraction,…

Computation and Language · Computer Science 2022-10-13 Platon Fedorov , Alexey Mironov , George Chernishev

In this paper, we present a study of neologisms and loan words frequently occurring in Facebook user posts. We have analyzed a dataset of several million publically available posts written during 2006-2013 by Russian-speaking Facebook…

Computation and Language · Computer Science 2018-04-18 Nikita Muravyev , Alexander Panchenko , Sergei Obiedkov

The paper presents RuBQ, the first Russian knowledge base question answering (KBQA) dataset. The high-quality dataset consists of 1,500 Russian questions of varying complexity, their English machine translations, SPARQL queries to Wikidata,…

Computation and Language · Computer Science 2021-10-14 Vladislav Korablinov , Pavel Braslavski

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of magnitude the largest…

Computation and Language · Computer Science 2022-04-13 Krishnapriya Vishnubhotla , Adam Hammond , Graeme Hirst

Citation analysis is one of the most frequently used methods in research evaluation. We are seeing significant growth in citation analysis through bibliometric metadata, primarily due to the availability of citation databases such as the…

Digital Libraries · Computer Science 2020-09-01 Sehrish Iqbal , Saeed-Ul Hassan , Naif Radi Aljohani , Salem Alelyani , Raheel Nawaz , Lutz Bornmann

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimodal data, which…

Computation and Language · Computer Science 2024-05-14 Alena Tsanda , Elena Bruches

The General QA field has been developing the methodology referencing the Stanford Question answering dataset (SQuAD) as the significant benchmark. However, compiling factual questions is accompanied by time- and labour-consuming annotation,…

Computation and Language · Computer Science 2022-11-11 Dina Pisarevskaya , Tatiana Shavrina

We describe a method of using statistically-collected Chinese character groups from a corpus to augment a Chinese dictionary. The method is particularly useful for extracting domain-specific and regional words not readily available in…

cmp-lg · Computer Science 2008-02-03 Pascale Fung , Dekai Wu

The evolution of vocabulary in academic publishing is characterized via keyword frequencies recorded the ISI Web of Science citations database. In four distinct case-studies, evolutionary analysis of keyword frequency change through time is…

Physics and Society · Physics 2009-02-18 R. Alexander Bentley

Quotes of public figures can mark turning points in history. A quote can explain its originator's actions, foreshadowing political or personal decisions and revealing character traits. Impactful quotes cross language barriers and influence…

Computation and Language · Computer Science 2022-07-21 Tin Kuculo , Simon Gottschalk , Elena Demidova

We study Twitter data from a dynamical systems perspective. In particular, we focus on the large set of data released by Twitter Inc. and asserted to represent a Russian influence operation. We propose a mathematical model to describe the…

Social and Information Networks · Computer Science 2020-01-28 Sarah Rajtmajer , Ashish Simhachalam , Thomas Zhao , Brady Bickel , Christopher Griffin

We present empirical data on misprints in citations to twelve high-profile papers. The great majority of misprints are identical to misprints in articles that earlier cited the same paper. The distribution of the numbers of misprint…

Physics and Society · Physics 2011-09-13 M. V. Simkin , V. P. Roychowdhury

To help individuals express themselves better, quotation recommendation is receiving growing attention. Nevertheless, most prior efforts focus on modeling quotations and queries separately and ignore the relationship between the quotations…

Computation and Language · Computer Science 2021-06-02 Lingzhi Wang , Xingshan Zeng , Kam-Fai Wong

In an article written five years ago [arXiv:0809.0522], we described a method for predicting which scientific papers will be highly cited in the future, even if they are currently not highly cited. Applying the method to real citation data…

Physics and Society · Physics 2014-02-06 M. E. J. Newman