中文
相关论文

相关论文: Large-scale diversity estimation through surname o…

200 篇论文

In this article, we develop and investigate a new classifier based on features extracted using spatial depth. Our construction is based on fitting a generalized additive model to the posterior probabilities of the different competing…

统计方法学 · 统计学 2015-04-16 Subhajit Dutta , Anil K. Ghosh

This study deals with a fairly simply formulated problem -- how to estimate the number of people bearing the same full name in a large population. Estimation of name popularity can leverage personal name matching in databases and be of…

数据库 · 计算机科学 2021-10-14 Ksenia Zhagorina , Pavel Braslavski , Vladimir Gusev

The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic…

Statistical consistency in phylogenetics has traditionally referred to the accuracy of estimating phylogenetic parameters for a fixed number of species as we increase the number of characters. However, as sequences are often of fixed length…

种群与进化 · 定量生物学 2010-04-09 Olivier Gascuel , Mike Steel

What makes some types of languages more probable than others? For instance, we know that almost all spoken languages contain the vowel phoneme /i/; why should that be? The field of linguistic typology seeks to answer these questions and,…

计算与语言 · 计算机科学 2018-07-10 Ryan Cotterell , Jason Eisner

Linguistic typology aims to capture structural and semantic variation across the world's languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that…

The recent availability of larger typological databases such as Haspelmath et al. (2005) has brought the linguistics community closer to having a solid, empirical foundation for making actual claims about prehistoric migrations, deep…

物理与社会 · 物理学 2007-05-23 Eric W. Holman , Christian Schulze , Dietrich Stauffer , Soren Wichmann

Diversity is an important property of datasets and sampling data for diversity is useful in dataset creation. Finding the optimally diverse sample is expensive, we therefore present a heuristic significantly increasing diversity relative to…

计算与语言 · 计算机科学 2025-01-15 Louis Estève , Manon Scholivet , Agata Savary

Genomes may be analyzed from an information viewpoint as very long strings, containing functional elements of variable length, which have been assembled by evolution. In this work an innovative information theory based algorithm is…

基因组学 · 定量生物学 2020-09-23 Vincenzo Bonnici , Giuditta Franco , Vincenzo Manca

Many questions that we have about the history and dynamics of organisms have a geographical component: How many are there, and where do they live? How do they move and interbreed across the landscape? How were they moving a thousand years…

种群与进化 · 定量生物学 2019-11-28 Gideon S. Bradburd , Peter L. Ralph

Traditionally, heritability has been estimated using family-based methods such as twin studies. Advancements in molecular genomics have facilitated the development of alternative methods that utilise large samples of unrelated or related…

The multivariate hypergeometric distribution describes sampling without replacement from a discrete population of elements divided into multiple categories. Addressing a gap in the literature, we tackle the challenge of estimating discrete…

机器学习 · 计算机科学 2024-06-11 Liam Hodgson , Danilo Bzdok

Phylogenetic diversity indices are commonly used to rank the elements in a collection of species or populations for conservation purposes. The derivation of these indices is typically based on some quantitative description of the…

种群与进化 · 定量生物学 2024-03-26 Vincent Moulton , Andreas Spillner , Kristina Wicke

Large language models (LLMs) are widely applied across diverse domains, raising concerns about their limitations and potential risks. In this study, we investigate two types of bias that LLMs may display: stereotype bias and deviation bias.…

计算与语言 · 计算机科学 2026-05-20 Daniel Wang , Eli Brignac , Minjia Mao , Xiao Fang

The number of extant individuals within a lineage, as exemplified by counts of species numbers across genera in a higher taxonomic category, is known to be a highly skewed distribution. Because the sublineages (such as genera in a clade)…

应用统计 · 统计学 2009-01-09 Panagis Moschopoulos , Max Shpak

We study in details the turnout rate statistics for 77 elections in 11 different countries. We show that the empirical results established in a previous paper for French elections appear to hold much more generally. We find in particular…

物理与社会 · 物理学 2015-06-03 Christian Borghesi , Jean-Claude Raynal , Jean-Philippe Bouchaud

The paper reviews the results obtained for spatial population models and the evolution of the genealogies of these populations during the last decade by the author and his coworkers. The focus is on their large scale behaviour and on the…

概率论 · 数学 2019-07-17 Andreas Greven

Goods, styles, ideologies are adopted by society through various mechanisms. In particular, adoption driven by innovation is extensively studied by marketing economics. Mathematical models are currently used to forecast the sales of…

物理与社会 · 物理学 2014-04-02 Baptiste Coulmont , Virginie Supervie , Romulus Breban

This paper presents exploratory techniques for multivariate data, many of them well known to French statisticians and ecologists, but few well understood in North American culture. We present the general framework of duality diagrams which…

统计方法学 · 统计学 2008-12-18 Susan Holmes

For the last few years, the amount of data has significantly increased in the companies. It is the reason why data analysis methods have to evolve to meet new demands. In this article, we introduce a practical analysis of a large database…

机器学习 · 统计学 2015-11-02 Romain Guigourès , Marc Boullé , Fabrice Rossi