English
Related papers

Related papers: Analysis of the quotation corpus of the Russian Wi…

200 papers

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-22 Gaspard Michel , Elena V. Epure , Christophe Cerisara

We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the…

Computation and Language · Computer Science 2013-09-10 Tomas Mikolov , Kai Chen , Greg Corrado , Jeffrey Dean

Rankings of scholarly journals based on citation data are often met with skepticism by the scientific community. Part of the skepticism is due to disparity between the common perception of journals' prestige and their ranking based on…

Applications · Statistics 2015-12-16 Cristiano Varin , Manuela Cattelan , David Firth

This paper describes the results of the first shared task on taxonomy enrichment for the Russian language. The participants were asked to extend an existing taxonomy with previously unseen words: for each new word their systems should…

Computation and Language · Computer Science 2020-05-25 Irina Nikishina , Varvara Logacheva , Alexander Panchenko , Natalia Loukachevitch

Natural language interfaces (NLIs) for data visualization are becoming increasingly popular both in academic research and in commercial software. Yet, there is a lack of empirical understanding of how people specify visualizations through…

Human-Computer Interaction · Computer Science 2021-10-05 Arjun Srinivasan , Nikhila Nyapathy , Bongshin Lee , Steven M. Drucker , John Stasko

The paper presents a study of methods for extracting information about dialogue participants and evaluating their performance in Russian. To train models for this task, the Multi-Session Chat dataset was translated into Russian using…

Computation and Language · Computer Science 2024-07-15 Konstantin Zaitsev

Computational Humor involves several tasks, such as humor recognition, humor generation, and humor scoring, for which it is useful to have human-curated data. In this work we present a corpus of 27,000 tweets written in Spanish and…

Computation and Language · Computer Science 2018-07-20 Santiago Castro , Luis Chiruzzo , Aiala Rosá , Diego Garat , Guillermo Moncecchi

The Classification Literature Automated Search Service, an annual bibliography based on citation of one or more of a set of around 80 book or journal publications, ran from 1972 to 2012. We analyze here the years 1994 to 2011. The…

Digital Libraries · Computer Science 2013-08-20 Fionn Murtagh , Michael J. Kurtz

Citation parsing is fundamental for search engines within academia and the protection of intellectual property. Meticulous extraction is further needed when evaluating the similarity of documents and calculating their citation impact.…

Digital Libraries · Computer Science 2018-05-23 Niall Martin Ryan

The paper introduces manually annotated test sets for the task of tracing diachronic (temporal) semantic shifts in Russian. The two test sets are complementary in that the first one covers comparatively strong semantic changes occurring to…

Computation and Language · Computer Science 2019-07-31 Vadim Fomin , Daria Bakshandaeva , Julia Rodina , Andrey Kutuzov

Publicly available data reveal long-term systematic features about citation statistics and how papers are referenced. The data also tell fascinating citation histories of individual articles.

Physics and Society · Physics 2009-11-11 S. Redner

The paper is a survey of notions and results related to classical and new generalizations of the notion of a periodic sequence. The topics related to almost periodicity in combinatorics on words, symbolic dynamics, expressibility in logical…

Discrete Mathematics · Computer Science 2015-05-13 An. A. Muchnik , Yu. L. Pritykin , A. L. Semenov

Here I present an investigation on the evolution and use of vocabulary in data science in the last 13 years. Based on a rigorous statistical analysis, a database with 12,787 documents containing the words "data science" in the title,…

Digital Libraries · Computer Science 2022-04-22 Igor Barahona

This paper surveys 60 English Machine Reading Comprehension datasets, with a view to providing a convenient resource for other researchers interested in this problem. We categorize the datasets according to their question and answer form…

Computation and Language · Computer Science 2021-10-11 Daria Dzendzik , Carl Vogel , Jennifer Foster

With social media datasets being increasingly shared by researchers, it also presents the caveat that those datasets are not always completely replicable. Having to adhere to requirements of platforms like Twitter, researchers cannot…

Digital Libraries · Computer Science 2018-03-08 Arkaitz Zubiaga

We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split…

Computation and Language · Computer Science 2019-09-27 Christina Niklaus , Andre Freitas , Siegfried Handschuh

We present NEWSROOM, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these…

Computation and Language · Computer Science 2020-05-19 Max Grusky , Mor Naaman , Yoav Artzi

The paper presents the results of the study of international and national cooperation of Ukrainian scientists from different scientific fields using citation analysis data from Scopus in the period of 2011-2015. The results show that during…

Digital Libraries · Computer Science 2018-07-30 Serhii Nazarovets

We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000…

Computation and Language · Computer Science 2019-08-21 Peter D. Turney , Saif M. Mohammad

Academic institutions, federal agencies, publishers, editors, authors, and librarians increasingly rely on citation analysis for making hiring, promotion, tenure, funding, and/or reviewer and journal evaluation and selection decisions. The…

Digital Libraries · Computer Science 2007-05-23 Lokman I. Meho , Kiduk Yang