Related papers: Zipf-Mandelbrot-Pareto model for co-authorship pop…
Publication statistics are ubiquitous in the ratings of scientific achievement, with citation counts and paper tallies factoring into an individual's consideration for postdoctoral positions, junior faculty, tenure, and even visa status for…
We have collected and cleaned two network data sets: Coauthorship and Citation networks for statisticians. The data sets are based on all research papers published in four of the top journals in statistics from $2003$ to the first half of…
Academic data sharing is a way for researchers to collaborate and thereby meet the needs of an increasingly complex research landscape. It enables researchers to verify results and to pursuit new research questions with "old" data. It is…
As science advances, the academic community has published millions of research papers. Researchers devote time and effort to search relevant manuscripts when writing a paper or simply to keep up with current research. In this paper, we…
Many complex systems are composed of disparate, interacting types of varying sizes: Species abundances in ecosystems, firm sizes in markets, city populations in countries, word counts in language, etc. A longstanding mystery of complex…
In the last years, researchers have realized the difficulties of fitting power-law distributions properly. These difficulties are higher in Zipf's systems, due to the discreteness of the variables and to the existence of two representations…
Scientific article recommender systems are playing an increasingly important role for researchers in retrieving scientific articles of interest in the coming era of big scholarly data. Most existing studies have designed unified methods for…
Ranking scientific authors is an important but challenging task, mostly due to the dynamic nature of the evolving scientific publications. The basic indicators of an author's productivity and impact are still the number of publications and…
It is widely recognized that collaboration between the public and private research sectors should be stimulated and supported, as a means of favoring innovation and regional development. This work takes a bibliometric approach, based on…
Academic citation and social attention measure different dimensions of the impact of research results. Both measures do not correlate with each other, and they are influenced by many factors. Among these factors are the field of research,…
Collaborations are an integral part of scientific research and publishing. In the past, access to large-scale corpora has limited the ways in which questions about collaborations could be investigated. However, with improvements in…
It turns out that some empirical facts in Big Data are the effects of properties of large numbers. Zipf's law 'noise' is an example of such an artefact. We expose several properties of the power law distributions and of similar distribution…
This paper explores various distributional aspects of random variables defined as the ratio of two independent positive random variables where one variable has an $\alpha$-stable law, for $0<\alpha<1$, and the other variable has the law…
Predicting the popularity of scientific publications has attracted many attentions from various disciplines. In this paper, we focus on the popularity prediction problem of scientific papers, and propose an age-based diffusion (AD) model to…
In the recent article Janosov, Battiston, & Sinatra report that they separated the inputs of talent and luck in creative careers. They build on the previous work of Sinatra et al which introduced the Q-model. Under the model the popularity…
High dimensional case control studies are ubiquitous in the biological sciences, particularly genomics. To maximise power while constraining cost and to minimise type-1 error rates, researchers typically seek to replicate findings in a…
A fair assignment of credit for multi-authored publications is a long-standing issue in scientometrics. In the calculation of the $h$-index, for instance, all co-authors receive equal credit for a given publication, independent of a given…
Here we sketch a new derivation of Zipf's law for word frequencies based on optimal coding. The structure of the derivation is reminiscent of Mandelbrot's random typing model but it has multiple advantages over random typing: (1) it starts…
This exploratory study examines how low-impact journals, defined through subject-normalized Eigenfactor percentiles, are associated with denser and more reciprocating patterns of author-to-author citations. Using Crossref records, we assign…
We propose measures of the impact of research that improve on existing ones such as counting of number of papers, citations and $h$-index. Since different papers and different fields have largely different average number of co-authors and…