中文
相关论文

相关论文: Analysis of the quotation corpus of the Russian Wi…

200 篇论文

Style is an important concept in today's challenges in natural language generating. After the success in the field of image style transfer, the task of text style transfer became actual and attractive. Researchers are also interested in the…

计算与语言 · 计算机科学 2023-06-06 Boris Orekhov

Dynamics of average length of words in Russian and English is analysed in the article. Words belonging to the diachronic text corpus Google Books Ngram and dated back to the last two centuries are studied. It was found out that average word…

计算与语言 · 计算机科学 2016-05-26 Vladimir V. Bochkarev , Anna V. Shevlyakova , Valery D. Solovyev

Citations are widely considered in scientists' evaluation. As such, scientists may be incentivized to inflate their citation counts. While previous literature has examined self-citations and citation cartels, it remains unclear whether…

计算工程、金融与科学 · 计算机科学 2024-02-08 Hazem Ibrahim , Fengyuan Liu , Yasir Zaki , Talal Rahwan

Recent advancements in Natural Language Processing (NLP) have fostered the development of Large Language Models (LLMs) that can solve an immense variety of tasks. One of the key aspects of their application is their ability to work with…

In this paper, we present a novel series of Russian information retrieval datasets constructed from the "Did you know..." section of Russian Wikipedia. Our datasets support a range of retrieval tasks, including fact-checking,…

信息检索 · 计算机科学 2025-11-10 Grigory Kovalev , Natalia Loukachevitch , Mikhail Tikhomirov , Olga Babina , Pavel Mamaev

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

计算与语言 · 计算机科学 2019-09-20 Alessia Battisti , Sarah Ebling

We present ToTTo, an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description.…

计算与语言 · 计算机科学 2020-10-07 Ankur P. Parikh , Xuezhi Wang , Sebastian Gehrmann , Manaal Faruqui , Bhuwan Dhingra , Diyi Yang , Dipanjan Das

Wikipedia can be edited by anyone and thus contains various quality sentences. Therefore, Wikipedia includes some poor-quality edits, which are often marked up by other editors. While editors' reviews enhance the credibility of Wikipedia,…

计算与语言 · 计算机科学 2024-01-02 Kenichiro Ando , Satoshi Sekine , Mamoru Komachi

We present RuSemShift, a large-scale manually annotated test set for the task of semantic change modeling in Russian for two long-term time period pairs: from the pre-Soviet through the Soviet times and from the Soviet through the…

计算与语言 · 计算机科学 2020-10-14 Julia Rodina , Andrey Kutuzov

Half a billion citation edges extracted from 100.7 million Ukrainian court decisions reveal that judicial citation structure encodes legal domain boundaries without supervision and predicts future legislative importance with near-perfect…

计算与语言 · 计算机科学 2026-05-18 Volodymyr Ovcharov

Interoperability is a feature required by the Semantic Web. It is provided by the ontology matching methods and algorithms. But now ontologies are presented not only in English, but in other languages as well. It is important to use an…

信息检索 · 计算机科学 2011-10-27 Feiyu Lin , Andrew Krizhanovsky

We study the statistics of citations from all Physical Review journals for the 110-year period 1893 until 2003. In addition to characterizing the citation distribution and identifying publications with the highest citation impact, we…

物理与社会 · 物理学 2007-05-23 S. Redner

Machine-generated citation sentences can aid automated scientific literature review and assist article writing. Current methods in generating citation text were limited to single citation generation using the citing document and a cited…

计算与语言 · 计算机科学 2021-12-10 Jia-Yan Wu , Alexander Te-Wei Shieh , Shih-Ju Hsu , Yun-Nung Chen

The algorithm of the creation texts parallel corpora was presented. The algorithm is based on the use of "key words" in text documents, and on the means of their automated translation. Key words were singled out by means of using Russian…

计算与语言 · 计算机科学 2008-07-03 D. V. Lande , V. V. Zhygalo

Citation numbers and other quantities derived from bibliographic databases are becoming standard tools for the assessment of productivity and impact of research activities. Though widely used, still their statistical properties have not…

物理与社会 · 物理学 2013-12-17 Filippo Radicchi , Claudio Castellano

Electronic dictionaries have largely replaced paper dictionaries and become central tools for L2 learners seeking to expand their vocabulary. Users often assume these resources are reliable and rarely question the validity of the…

计算与语言 · 计算机科学 2025-08-18 Shiyang Zhang , Fanfei Meng , Xi Wang , Lan Li

Citation information in scholarly data is an important source of insight into the reception of publications and the scholarly discourse. Outcomes of citation analyses and the applicability of citation based machine learning approaches…

数字图书馆 · 计算机科学 2022-01-12 Tarek Saier , Michael Färber , Tornike Tsereteli

The article describes the original method of creating a dictionary of abbreviations based on the Google Books Ngram Corpus. The dictionary of abbreviations is designed for Russian, yet as its methodology is universal it can be applied to…

计算与语言 · 计算机科学 2014-10-07 Valery D. Solovyev , Vladimir V. Bochkarev

In this paper, we introduce the Dialogue Evaluation shared task on extraction of structured opinions from Russian news texts. The task of the contest is to extract opinion tuples for a given sentence; the tuples are composed of a sentiment…

Books have been widely used to share information and contribute to human knowledge. However, the quantitative use of books as a method of scholarly communication is relatively unexamined compared to journal articles and conference papers.…

数字图书馆 · 计算机科学 2019-12-09 Yongjun Zhu , Erjia Yan , Silvio Peroni , Chao Che