English
Related papers

Related papers: Analysis of the quotation corpus of the Russian Wi…

200 papers

In this paper, we present Russian language datasets in the digital humanities domain for the evaluation of word embedding techniques or similar language modeling and feature learning algorithms. The datasets are split into two task types,…

Computation and Language · Computer Science 2019-03-22 Gerhard Wohlgenannt , Artemii Babushkin , Denis Romashov , Igor Ukrainets , Anton Maskaykin , Ilya Shutov

Trends are analysed in the annual number of documents published by Russian institutions and indexed in Scopus and Web of Science, giving special attention to the time period starting in the year 2013 in which the Project 5-100 was launched…

Digital Libraries · Computer Science 2018-05-08 Henk F. Moed , Valentina Markusova , Mark Akoev

Information Extraction is a well-researched area of Natural Language Processing with applications in web search and question answering concerned with identifying entities and relationships between them as expressed in a given context,…

Information Retrieval · Computer Science 2020-11-17 Erin Macdonald , Denilson Barbosa

Extracting who says what to whom is a crucial part in analyzing human communication in today's abundance of data such as online news articles. Yet, the lack of annotated data for this task in German news articles severely limits the quality…

Computation and Language · Computer Science 2024-04-26 Fynn Petersen-Frey , Chris Biemann

Proverbs are an essential component of language and culture, and though much attention has been paid to their history and currency, there has been comparatively little quantitative work on changes in the frequency with which they are used…

Computation and Language · Computer Science 2021-07-13 E. Davis , C. M. Danforth , W. Mieder , P. S. Dodds

Simile is a figure of speech that compares two things through the use of connection words, but where comparison is not intended to be taken literally. They are often used in everyday communication, but they are also a part of linguistic…

Computation and Language · Computer Science 2018-11-27 Nikola Milosevic , Goran Nenadic

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

Computation and Language · Computer Science 2020-04-03 Jonathan Dunn

While Large language models (LLMs) have become excellent writing assistants, they still struggle with quotation generation. This is because they either hallucinate when providing factual quotations or fail to provide quotes that exceed…

Computation and Language · Computer Science 2025-02-21 Jin Xiao , Bowei Zhang , Qianyu He , Jiaqing Liang , Feng Wei , Jinglei Chen , Zujie Liang , Deqing Yang , Yanghua Xiao

In this paper, we present a distributional word embedding model trained on one of the largest available Russian corpora: Araneum Russicum Maximum (over 10 billion words crawled from the web). We compare this model to the model trained on…

Computation and Language · Computer Science 2018-01-22 Andrey Kutuzov , Maria Kunilovskaya

We present the shared task on artificial text detection in Russian, which is organized as a part of the Dialogue Evaluation initiative, held in 2022. The shared task dataset includes texts from 14 text generators, i.e., one human writer and…

Our current knowledge of scholarly plagiarism is largely based on the similarity between full text research articles. In this paper, we propose an innovative and novel conceptualization of scholarly plagiarism in the form of reuse of…

Digital Libraries · Computer Science 2017-05-09 Mayank Singh , Abhishek Niranjan , Divyansh Gupta , Nikhil Angad Bakshi , Animesh Mukherjee , Pawan Goyal

Linguistic acceptability (LA) attracts the attention of the research community due to its many uses, such as testing the grammatical knowledge of language models and filtering implausible texts with acceptability classifiers. However, the…

Computation and Language · Computer Science 2023-10-04 Vladislav Mikhailov , Tatiana Shamardina , Max Ryabinin , Alena Pestova , Ivan Smurov , Ekaterina Artemova

With over 20 million records, the ADS citation database is regularly used by researchers and librarians to measure the scientific impact of individuals, groups, and institutions. In addition to the traditional sources of citations, the ADS…

Generative poetry systems require effective tools for data engineering and automatic evaluation, particularly to assess how well a poem adheres to versification rules, such as the correct alternation of stressed and unstressed syllables and…

Computation and Language · Computer Science 2025-10-21 Ilya Koziev

It is very common to use quotations (quotes) to make our writings more elegant or convincing. To help people find appropriate quotes efficiently, the task of quote recommendation is presented, aiming to recommend quotes that fit the current…

Computation and Language · Computer Science 2022-03-15 Fanchao Qi , Yanhui Yang , Jing Yi , Zhili Cheng , Zhiyuan Liu , Maosong Sun

Russian Internet Trolls use fake personas to spread disinformation through multiple social media streams. Given the increased frequency of this threat across social media platforms, understanding those operations is paramount in combating…

Social and Information Networks · Computer Science 2024-09-16 Sachith Dassanayaka , Ori Swed , Dimitri Volchenkov

Current approaches to automatic summarization of scientific papers generate informative summaries in the form of abstracts. However, abstracts are not intended to show the relationship between a paper and the references cited in it. We…

Computation and Language · Computer Science 2023-11-14 Shahbaz Syed , Ahmad Dawar Hakimi , Khalid Al-Khatib , Martin Potthast

Similes are natural language expressions used to compare unlikely things, where the comparison is not taken literally. They are often used in everyday communication and are an important part of cultural heritage. Having an up-to-date corpus…

Computation and Language · Computer Science 2016-05-23 Nikola Milosevic , Goran Nenadic

This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks. We present a new Multi-Topic…

Computation and Language · Computer Science 2022-11-09 Yao Dou , Chao Jiang , Wei Xu

Statistical analysis of repeat misprints in scientific citations leads to the conclusion that about 80% of scientific citations are copied from the lists of references used in othe papers. Based on this finding a mathematical theory of…

Statistics Theory · Mathematics 2007-06-13 M. V. Simkin , V. P. Roychowdhury