中文
相关论文

相关论文: Plagiarism Detection in arXiv

200 篇论文

Hosting over 10 million of software projects, GitHub is one of the most important data sources to study behavior of developers and software projects. However, with the increase of the size of open source datasets, the potential threats to…

软件工程 · 计算机科学 2018-05-09 Can Cheng , Bing Li , Zengyang Li , Peng Liang

Authorship ethics is a central topic of discussion in research ethics fora. There are various guidelines for authorship (i.e., naming and order). It is not easy to decide the authorship in the presence of varying authorship guidelines. This…

软件工程 · 计算机科学 2021-03-29 Nasir Mehmood Minhas

The rate at which scholarly literature is being produced has been increasing at approximately 3.5 percent per year for decades. This means that during a typical 40 year career the amount of new literature produced each year increases by a…

信息检索 · 计算机科学 2022-12-21 Michael J. Kurtz , Edwin A. Henneken

We address the problem of counting the number of strings in a collection where a given pattern appears, which has applications in information retrieval and data mining. Existing solutions are in a theoretical stage. We implement these…

数据结构与算法 · 计算机科学 2015-10-02 Travis Gagie , Aleksi Hartikainen , Juha Kärkkäinen , Gonzalo Navarro , Simon J. Puglisi , Jouni Sirén

Handwritten document analysis is an area of forensic science, with the goal of establishing authorship of documents through examination of inherent characteristics. Law enforcement agencies use standard protocols based on manual processing…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Eleonora Breci , Luca Guarnera , Sebastiano Battiato

The article presents a system for testing the independence of solutions to algorithmic problems sent by students as part of the student programming competition. First, the context was discussed, as well as the need to organize programming…

软件工程 · 计算机科学 2019-12-18 Zenon Gniazdowski , Maciej Boniecki

The main objective of this paper is to identify the set of highly-cited documents in Google Scholar and to define their core characteristics (document types, language, free availability, source providers, and number of versions), under the…

数字图书馆 · 计算机科学 2016-12-28 Alberto Martin-Martin , Enrique Orduna-Malea , Juan M. Ayllon , Emilio Delgado Lopez-Cozar

Plagiarism detection is a growing need among educational institutions and solutions for different purposes exist. An important field in this direction is detecting cases of source-code plagiarism. In this paper, we present the tool Kato for…

计算机科学中的逻辑 · 计算机科学 2010-07-29 Johannes Oetsch , Jörg Pührer , Martin Schwengerer , Hans Tompits

Texts exhibit considerable stylistic variation. This paper reports an experiment where a corpus of documents (N= 75 000) is analyzed using various simple stylistic metrics. A subset (n = 1000) of the corpus has been previously assessed to…

cmp-lg · 计算机科学 2008-02-03 Jussi Karlgren

The amount of scholarly data has been increasing dramatically over the last years. For newcomers to a particular science domain (e.g., IR, physics, NLP) it is often difficult to spot larger trends and to position the latest research in the…

数字图书馆 · 计算机科学 2021-12-08 Naman Paharia , Muhammad Syafiq Mohd Pozi , Adam Jatowt

Code analyzers such as Error Prone and FindBugs detect code patterns symptomatic of bugs, performance issues, or bad style. These tools express patterns as quick fixes that detect and rewrite unwanted code. However, it is difficult to come…

软件工程 · 计算机科学 2018-09-11 Reudismam Rolim , Gustavo Soares , Rohit Gheyi , Titus Barik , Loris D'Antoni

Automatic sarcasm detection is the task of predicting sarcasm in text. This is a crucial step to sentiment analysis, considering prevalence and challenges of sarcasm in sentiment-bearing text. Beginning with an approach that used…

计算与语言 · 计算机科学 2016-09-22 Aditya Joshi , Pushpak Bhattacharyya , Mark James Carman

The proliferation of data and text documents such as articles, web pages, books, social network posts, etc. on the Internet has created a fundamental challenge in various fields of text processing under the title of "automatic text…

人工智能 · 计算机科学 2023-03-15 Kazem Taghandiki , Mohammad Hassan Ahmadi , Elnaz Rezaei Ehsan

Sarcasm detection is an important task in affective computing, requiring large amounts of labeled data. We introduce reactive supervision, a novel data collection method that utilizes the dynamics of online conversations to overcome the…

计算与语言 · 计算机科学 2023-09-07 Boaz Shmueli , Lun-Wei Ku , Soumya Ray

Usability issues can hinder the effective use of software. Therefore, various techniques are deployed to diagnose and mitigate them. However, these techniques are costly and time-consuming, particularly in iterative design and development.…

人机交互 · 计算机科学 2025-04-03 Eduard Kuric , Peter Demcak , Matus Krajcovic , Jan Lang

Preprints play an increasingly critical role in academic communities. There are many reasons driving researchers to post their manuscripts to preprint servers before formal submission to journals or conferences, but the use of preprints has…

数字图书馆 · 计算机科学 2023-08-04 Jialiang Lin , Yao Yu , Yu Zhou , Zhiyang Zhou , Xiaodong Shi

We explore the degree to which papers prepublished on arXiv garner more citations, in an attempt to paint a sharper picture of fairness issues related to prepublishing. A paper's citation count is estimated using a negative-binomial…

数字图书馆 · 计算机科学 2018-05-15 Sergey Feldman , Kyle Lo , Waleed Ammar

The scientific image integrity area presents a challenging research bottleneck, the lack of available datasets to design and evaluate forensic techniques. Its data sensitivity creates a legal hurdle that prevents one to rely on real…

计算机视觉与模式识别 · 计算机科学 2024-09-30 João P. Cardenuto , Anderson Rocha

Scientific writing is an iterative process that generates rich revision traces, yet publicly available resources typically expose only final or near-final versions of papers. This limits empirical study of revision behaviour and evaluation…

计算与语言 · 计算机科学 2026-03-31 Léane Jourdan , Julien Aubert-Béduchaud , Yannis Chupin , Marah Baccari , Florian Boudin

Plagiarism detection is one of the most researched areas among the Natural Language Processing(NLP) community. A good plagiarism detection covers all the NLP methods including semantics, named entities, paraphrases etc. and produces…

计算与语言 · 计算机科学 2023-08-25 Sagar Kulkarni , Sharvari Govilkar , Dhiraj Amin