中文
相关论文

相关论文: Scientific evaluation of Charles Dickens

200 篇论文

A number of applications (e.g., AI bot tournaments, sports, peer grading, crowdsourcing) use pairwise comparison data and the Bradley-Terry-Luce (BTL) model to evaluate a given collection of items (e.g., bots, teams, students, search…

机器学习 · 计算机科学 2019-06-12 Jingyan Wang , Nihar B. Shah , R. Ravi

We use an information-theoretic measure of linguistic similarity to investigate the organization and evolution of scientific fields. An analysis of almost 20M papers from the past three decades reveals that the linguistic similarity is…

数字图书馆 · 计算机科学 2018-01-30 Laercio Dias , Martin Gerlach , Joachim Scharloth , Eduardo G. Altmann

Automatic readability assessment plays a key role in ensuring effective and accessible written communication. Despite significant progress, the field is hindered by inconsistent definitions of readability and measurements that rely on…

计算与语言 · 计算机科学 2025-10-20 Catarina G Belem , Parker Glenn , Alfy Samuel , Anoop Kumar , Daben Liu

While living in different historical era, Charles Darwin (1809-1882) and Albert Einstein (1879-1955) were both prolific correspondents: Darwin sent (received) at least 7,591 (6,530) letters during his lifetime while Einstein sent (received)…

物理与社会 · 物理学 2009-11-11 J. G. Oliveira , A. -L. Barabási

With the help of online tools, unscrupulous authors can today generate a pseudo-scientific article and attempt to publish it. Some of these tools work by replacing or paraphrasing existing texts to produce new content, but they have a…

计算与语言 · 计算机科学 2022-10-25 Puthineath Lay , Martin Lentschat , Cyril Labbé

The James-Stein estimator's dominance over maximum likelihood in terms of mean square error (MSE) has been one of the most celebrated results in modern statistics, suggesting that biased estimators can systematically outperform unbiased…

统计理论 · 数学 2025-08-12 Paul W. Vos

The accuracy of published medical research is critical both for scientists, physicians and patients who rely on these results. But the fundamental belief in the medical literature was called into serious question by a paper suggesting most…

应用统计 · 统计学 2013-01-17 Leah R. Jager , Jeffrey T. Leek

In this paper we build on earlier observations and theory regarding word length frequency and sequential distribution to develop a mathematical characterization of some of the language features distinguishing isometrically lineated text…

cmp-lg · 计算机科学 2007-05-23 Hideaki Aoyama , John Constable

The problem of online threats and abuse could potentially be mitigated with a computational approach, where sources of abuse are better understood or identified through author profiling. However, abusive language constitutes a specific…

计算与语言 · 计算机科学 2020-09-04 Isabelle van der Vegt , Bennett Kleinberg , Paul Gill

The missionary zeal of many Bayesians of old has been matched, in the other direction, by a view among some theoreticians that Bayesian methods are absurd-not merely misguided but obviously wrong in principle. We consider several examples,…

统计理论 · 数学 2012-06-12 Andrew Gelman , Christian P. Robert

Recent work raises concerns about the use of standard splits to compare natural language processing models. We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to…

计算与语言 · 计算机科学 2020-10-08 Piotr Szymański , Kyle Gorman

For testing goodness of fit it is very popular to use either the chi square statistic or G statistics (information divergence). Asymptotically both are chi square distributed so an obvious question is which of the two statistics that has a…

统计理论 · 数学 2012-06-19 Peter Harremoës , Gábor Tusnády

This paper is concerned with a comparison of van der Waerden's and Wilcoxon's scores.

统计理论 · 数学 2015-03-19 Nadir Maaroufi , Camille Sabbah , Yvik Swan , Thomas Verdebout

Using data from computer databases of scientific papers in physics, biomedical research, and computer science, we have constructed networks of collaboration between scientists in each of these disciplines. In these networks two scientists…

统计力学 · 物理学 2009-10-31 M. E. J. Newman

Can humans tell whether a news article was written by a person or a large language model (LLM)? We investigate this question using JudgeGPT, a study platform that independently measures source attribution (human vs. machine) and…

计算机与社会 · 计算机科学 2026-04-07 Alexander Loth , Martin Kappes , Marc-Oliver Pahl

How do two distributions of texts differ? Humans are slow at answering this, since discovering patterns might require tediously reading through hundreds of samples. We propose to automatically summarize the differences by "learning a…

计算与语言 · 计算机科学 2022-05-19 Ruiqi Zhong , Charlie Snell , Dan Klein , Jacob Steinhardt

Recent advancements in neural language modelling make it possible to rapidly generate vast amounts of human-sounding text. The capabilities of humans and automatic discriminators to detect machine-generated text have been a large source of…

计算与语言 · 计算机科学 2020-05-11 Daphne Ippolito , Daniel Duckworth , Chris Callison-Burch , Douglas Eck

Story-writing is a fundamental aspect of human imagination, relying heavily on creativity to produce narratives that are novel, effective, and surprising. While large language models (LLMs) have demonstrated the ability to generate…

计算与语言 · 计算机科学 2025-05-13 Mete Ismayilzada , Claire Stevenson , Lonneke van der Plas

This study presents a theoretical analysis on the efficiency of interleaving, an efficient online evaluation method for rankings. Although interleaving has already been applied to production systems, the source of its high efficiency has…

信息检索 · 计算机科学 2023-06-21 Kojiro Iizuka , Hajime Morita , Makoto P. Kato

This is a report about the use and misuse of citation data in the assessment of scientific research. The idea that research assessment must be done using ``simple and objective'' methods is increasingly prevalent today. The ``simple and…

统计方法学 · 统计学 2009-10-20 Robert Adler , John Ewing , Peter Taylor
‹ 上一页 1 8 9 10 下一页 ›