中文
相关论文

相关论文: Analysis of the quotation corpus of the Russian Wi…

200 篇论文

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works,…

音频与语音处理 · 电气工程与系统科学 2026-04-22 Gaspard Michel , Elena V. Epure , Christophe Cerisara

We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the…

计算与语言 · 计算机科学 2013-09-10 Tomas Mikolov , Kai Chen , Greg Corrado , Jeffrey Dean

Rankings of scholarly journals based on citation data are often met with skepticism by the scientific community. Part of the skepticism is due to disparity between the common perception of journals' prestige and their ranking based on…

应用统计 · 统计学 2015-12-16 Cristiano Varin , Manuela Cattelan , David Firth

This paper describes the results of the first shared task on taxonomy enrichment for the Russian language. The participants were asked to extend an existing taxonomy with previously unseen words: for each new word their systems should…

计算与语言 · 计算机科学 2020-05-25 Irina Nikishina , Varvara Logacheva , Alexander Panchenko , Natalia Loukachevitch

Natural language interfaces (NLIs) for data visualization are becoming increasingly popular both in academic research and in commercial software. Yet, there is a lack of empirical understanding of how people specify visualizations through…

人机交互 · 计算机科学 2021-10-05 Arjun Srinivasan , Nikhila Nyapathy , Bongshin Lee , Steven M. Drucker , John Stasko

The paper presents a study of methods for extracting information about dialogue participants and evaluating their performance in Russian. To train models for this task, the Multi-Session Chat dataset was translated into Russian using…

计算与语言 · 计算机科学 2024-07-15 Konstantin Zaitsev

Computational Humor involves several tasks, such as humor recognition, humor generation, and humor scoring, for which it is useful to have human-curated data. In this work we present a corpus of 27,000 tweets written in Spanish and…

计算与语言 · 计算机科学 2018-07-20 Santiago Castro , Luis Chiruzzo , Aiala Rosá , Diego Garat , Guillermo Moncecchi

The Classification Literature Automated Search Service, an annual bibliography based on citation of one or more of a set of around 80 book or journal publications, ran from 1972 to 2012. We analyze here the years 1994 to 2011. The…

数字图书馆 · 计算机科学 2013-08-20 Fionn Murtagh , Michael J. Kurtz

Citation parsing is fundamental for search engines within academia and the protection of intellectual property. Meticulous extraction is further needed when evaluating the similarity of documents and calculating their citation impact.…

数字图书馆 · 计算机科学 2018-05-23 Niall Martin Ryan

The paper introduces manually annotated test sets for the task of tracing diachronic (temporal) semantic shifts in Russian. The two test sets are complementary in that the first one covers comparatively strong semantic changes occurring to…

计算与语言 · 计算机科学 2019-07-31 Vadim Fomin , Daria Bakshandaeva , Julia Rodina , Andrey Kutuzov

Publicly available data reveal long-term systematic features about citation statistics and how papers are referenced. The data also tell fascinating citation histories of individual articles.

物理与社会 · 物理学 2009-11-11 S. Redner

The paper is a survey of notions and results related to classical and new generalizations of the notion of a periodic sequence. The topics related to almost periodicity in combinatorics on words, symbolic dynamics, expressibility in logical…

离散数学 · 计算机科学 2015-05-13 An. A. Muchnik , Yu. L. Pritykin , A. L. Semenov

Here I present an investigation on the evolution and use of vocabulary in data science in the last 13 years. Based on a rigorous statistical analysis, a database with 12,787 documents containing the words "data science" in the title,…

数字图书馆 · 计算机科学 2022-04-22 Igor Barahona

This paper surveys 60 English Machine Reading Comprehension datasets, with a view to providing a convenient resource for other researchers interested in this problem. We categorize the datasets according to their question and answer form…

计算与语言 · 计算机科学 2021-10-11 Daria Dzendzik , Carl Vogel , Jennifer Foster

With social media datasets being increasingly shared by researchers, it also presents the caveat that those datasets are not always completely replicable. Having to adhere to requirements of platforms like Twitter, researchers cannot…

数字图书馆 · 计算机科学 2018-03-08 Arkaitz Zubiaga

We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split…

计算与语言 · 计算机科学 2019-09-27 Christina Niklaus , Andre Freitas , Siegfried Handschuh

We present NEWSROOM, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these…

计算与语言 · 计算机科学 2020-05-19 Max Grusky , Mor Naaman , Yoav Artzi

The paper presents the results of the study of international and national cooperation of Ukrainian scientists from different scientific fields using citation analysis data from Scopus in the period of 2011-2015. The results show that during…

数字图书馆 · 计算机科学 2018-07-30 Serhii Nazarovets

We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000…

计算与语言 · 计算机科学 2019-08-21 Peter D. Turney , Saif M. Mohammad

Academic institutions, federal agencies, publishers, editors, authors, and librarians increasingly rely on citation analysis for making hiring, promotion, tenure, funding, and/or reviewer and journal evaluation and selection decisions. The…

数字图书馆 · 计算机科学 2007-05-23 Lokman I. Meho , Kiduk Yang