Related papers: Sample Size for Concurrent Species Detection in a …
Natural ecosystems are characterized by striking diversity of form and functions and yet exhibit deep symmetries emerging across scales of space, time and organizational complexity. Species-area relationships and species-abundance…
Exploratory data analysis is crucial for developing and understanding classification models from high-dimensional datasets. We explore the utility of a new unsupervised tree ensemble called uncharted forest for visualizing class…
The difficulty of monitoring biodiversity at fine scales and over large areas limits ecological knowledge and conservation efforts. To fill this gap, Species Distribution Models (SDMs) predict species across space from spatially explicit…
The multispecies coalescent process models the genealogical relationships of genes sampled from several species, enabling useful predictions about phenomena such as the discordance between the gene tree and the species phylogeny due to…
Datasets may contain observations with multiple labels. If the labels are not mutually exclusive, and if the labels vary greatly in frequency, obtaining a sample that includes sufficient observations with scarcer labels to make inferences…
We investigate number count statistics as measures for transition to homogeneity of the matter distribution in the Universe and analyse how such statistics might be `dressed' by the assumed survey selection function. Since the estimated…
The twin crises of climate change and biodiversity loss define a strong need for functional diversity monitoring. While the availability of high-quality ecological monitoring data is increasing, the quantification of functional diversity so…
The wealth of data being gathered about humans and their surroundings drives new machine learning applications in various fields. Consequently, more and more often, classifiers are trained using not only numerical data but also complex data…
We propose a model for evolution aiming to reproduce statistical features of fossil data, in particular the distributions of extinction events, the distribution of species per genus and the distribution of lifetimes, all of which are known…
The distributions of species lifetimes and species in space are related, since species with good local survival chances have more time to colonize new habitats and species inhabiting large areas have higher chances to survive local…
Recent methodological advances are enabling better examination of speciation and extinction processes and patterns. A major open question is the origin of large discrepancies in species number between groups of the same age. Existing…
Measures of biodiversity change such as the Living Planet Index describe proportional change in the abundance of a typical species, which can be thought of as change in the size of a community. Here, I discuss the orthogonal concept of…
In this work we study the set size distribution estimation problem, where elements are randomly sampled from a collection of non-overlapping sets and we seek to recover the original set size distribution from the samples. This problem has…
Studies on distribution, abundance and diversity of species revealed fascinating universalities in macroecology. Many of these patterns, like the species-area and range-abundance relationship or the year-to-year fluctuations in population…
High dimensional and heterogeneous count data are collected in various applied fields. In this paper, we look closely at high-resolution sequencing data on the microbiome, which have enabled researchers to study the genomes of entire…
When analyzing communities of microorganisms from their sequenced DNA, an important task is taxonomic profiling: enumerating the presence and relative abundance of all organisms, or merely of all taxa, contained in the sample. This task can…
We derive a linear recursion relation for the species abundance distribution in a statistical model of ecology and demonstrate the existence of a scaling solution.
In an effort to catalog insect biodiversity, we propose a new large dataset of hand-labelled insect images, the BIOSCAN-Insect Dataset. Each record is taxonomically classified by an expert, and also has associated genetic information…
Learning joint probability distributions on n random variables requires exponential sample size in the generic case. Here we consider the case that a temporal (or causal) order of the variables is known and that the (unknown) graph of…
A large number of models of the species abundance distribution (SAD) have been proposed, many of which are generically similar to the log-normal distribution, from which they are often indistinguishable when describing a given data set.…