中文
相关论文

相关论文: Estimation of English and non-English Language Use…

200 篇论文

With language models becoming increasingly ubiquitous, it has become essential to address their inequitable treatment of diverse demographic groups and factors. Most research on evaluating and mitigating fairness harms has been concentrated…

计算与语言 · 计算机科学 2023-03-01 Krithika Ramesh , Sunayana Sitaram , Monojit Choudhury

We present a theory for the growth dynamics of the World Wide Web that takes into account the wide range of stochastic growth rates in the number of pages per site, as well as the fact that new sites are created at different times. This…

统计力学 · 物理学 2007-05-23 Bernardo A. Huberman , Lada A. Adamic

English has long been assumed the $\textit{lingua franca}$ of scientific research, and this notion is reflected in the natural language processing (NLP) research involving scientific document representation. In this position piece, we…

计算与语言 · 计算机科学 2024-03-28 Abteen Ebrahimi , Kenneth Church

In many real growing networks the mean number of connections per vertex increases with time. The Internet, the Word Wide Web, collaboration networks, and many others display this behavior. Such a growth can be called {\em accelerated}. We…

统计力学 · 物理学 2007-05-23 S. N. Dorogovtsev , J. F. F. Mendes

This paper presents the first attempt, up to our knowledge, to classify English writing styles on this scale with the challenge of classifying day to day language written by writers with different backgrounds covering various areas of…

计算与语言 · 计算机科学 2017-04-26 Yanging Chen , Rami Al-Rfou' , Yejin Choi

English Wikipedia has long been an important data source for much research and natural language machine learning modeling. The growth of non-English language editions of Wikipedia, greater computational resources, and calls for equity in…

计算机与社会 · 计算机科学 2022-04-07 Isaac Johnson , Emily Lescak

As one of the Web's primary multilingual knowledge sources, Wikipedia is read by millions of people across the globe every day. Despite this global readership, little is known about why users read Wikipedia's various language editions. To…

计算机与社会 · 计算机科学 2018-12-04 Florian Lemmerich , Diego Sáez-Trumper , Robert West , Leila Zia

Research in question answering datasets and models has gained a lot of attention in the research community. Many of them release their own question answering datasets as well as the models. There is tremendous progress that we have seen in…

计算与语言 · 计算机科学 2021-12-28 Andreas Chandra , Affandy Fahrizain , Ibrahim , Simon Willyanto Laufried

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

There is a great deal of work in cognitive psychology, linguistics, and computer science, about using word (or phrase) frequencies in context in text corpora to develop measures for word similarity or word association, going back to at…

计算与语言 · 计算机科学 2009-05-26 Rudi L. Cilibrasi , Paul M. B. Vitanyi

Using human evaluation of 100,000 words spread across 24 corpora in 10 languages diverse in origin and culture, we present evidence of a deep imprint of human sociality in language, observing that (1) the words of natural human language…

Natural language processing (NLP) systems have become a central technology in communication, education, medicine, artificial intelligence, and many other domains of research and development. While the performance of NLP methods has grown…

计算与语言 · 计算机科学 2021-10-14 Damián Blasi , Antonios Anastasopoulos , Graham Neubig

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

计算与语言 · 计算机科学 2015-03-05 Marcelo A Montemurro , Damián H Zanette

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…

计算与语言 · 计算机科学 2015-05-15 Germinal Cocho , Jorge Flores , Carlos Gershenson , Carlos Pineda , Sergio Sánchez

We analyze the rank-frequency distributions of words in selected English and Polish texts. We compare scaling properties of these distributions in both languages. We also study a few small corpora of Polish literary texts and find that for…

计算与语言 · 计算机科学 2015-11-11 Stanislaw Drozdz , Jaroslaw Kwapien , Adam Orczyk

We introduce the notion of the 'meaning bound' of a word with respect to another word by making use of the World-Wide Web as a conceptual environment for meaning. The meaning of a word with respect to another word is established by…

人工智能 · 计算机科学 2012-03-28 Diederik Aerts

Of basic interest is the quantification of the long term growth of a language's lexicon as it develops to more completely cover both a culture's communication requirements and knowledge space. Here, we explore the usage dynamics of words in…

计算与语言 · 计算机科学 2017-03-27 Eitan Adam Pechenick , Christopher M. Danforth , Peter Sheridan Dodds

The distribution of numbers in human documents is determined by a variety of diverse natural and human factors, whose relative significance can be evaluated by studying the numbers' frequency of occurrence. Although it has been studied…

物理与社会 · 物理学 2009-11-11 S. N. Dorogovtsev , J. F. F. Mendes , J. G. Oliveira

As the World Wide Web is growing rapidly, it is getting increasingly challenging to gather representative information about it. Instead of crawling the web exhaustively one has to resort to other techniques like sampling to determine the…

数据结构与算法 · 计算机科学 2009-02-11 Eda Baykan , Monika Henzinger , Stefan F. Keller , Sebastian De Castelberg , Markus Kinzler