English
Related papers

Related papers: Personal Names Popularity Estimation and its Appli…

200 papers

We consider an extension of the {\em popular matching} problem in this paper. The input to the popular matching problem is a bipartite graph G = (A U B,E), where A is a set of people, B is a set of items, and each person a belonging to A…

Data Structures and Algorithms · Computer Science 2010-09-15 Telikepalli Kavitha , Meghana Nasre , Prajakta Nimbhorkar

We introduce a new metric for measuring how well a model personalizes to a user's specific preferences. We define personalization as a weighting between performance on user specific data and performance on a more general global dataset that…

Machine Learning · Computer Science 2021-04-26 Reuben Brasher , Nat Roth , Justin Wagle

The frequency distribution of personal given names offers important evidence about the information economy. This paper presents data on the popularity of the most frequent personal given names (first names) in England and Wales over the…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Douglas A. Galbi

Matching households and individuals across different databases poses challenges due to the lack of unique identifiers, typographical errors, and changes in attributes over time. Record linkage tools play a crucial role in overcoming these…

Applications · Statistics 2024-04-09 Thais Pacheco Menezes , Thomas Brendan Murphy , Michael Fop

A new approach to estimate population size based on a stratified link-tracing sampling design is presented. The method extends on the Frank and Snijders (1994) approach by allowing for heterogeneity in the initial sample selection…

Methodology · Statistics 2017-09-25 Kyle Vincent

A person's gender is a crucial piece of information when performing research across a wide range of scientific disciplines, such as medicine, sociology, political science, and economics, to name a few. However, in increasing instances,…

Computation and Language · Computer Science 2023-08-25 Kriste Krstovski , Yao Lu , Ye Xu

We study a problem of quick detection of top-k Personalized PageRank lists. This problem has a number of important applications such as finding local cuts in large graphs, estimation of similarity distance and name disambiguation. In…

Networking and Internet Architecture · Computer Science 2010-08-24 Konstantin Avrachenkov , Nelly Litvak , Danil A. Nemirovsky , Elena Smirnova , Marina Sokol

There has been substantial recent interest in record linkage, attempting to group the records pertaining to the same entities from a large database lacking unique identifiers. This can be viewed as a type of "microclustering," with few…

Statistics Theory · Mathematics 2017-03-16 James E. Johndrow , Kristian Lum , David B. Dunson

Nine popular clustering methods are applied to 42 real data sets. The aim is to give a detailed characterisation of the methods by means of several cluster validation indexes that measure various individual aspects of the resulting clusters…

Methodology · Statistics 2021-02-09 Christian Hennig

Music prediction tasks range from predicting tags given a song or clip of audio, predicting the name of the artist, or predicting related songs given a song, clip, artist name or tag. That is, we are interested in every semantic…

Machine Learning · Computer Science 2015-03-19 Jason Weston , Samy Bengio , Philippe Hamel

The problem of collecting reliable estimates of occurrence of entities on the open web forms the premise for this report. The models learned for tagging entities cannot be expected to perform well when deployed on the web. This is owing to…

Computation and Language · Computer Science 2016-05-17 Aman Madaan , Sunita Sarawagi

Matching is an important tool in causal inference. The method provides a conceptually straightforward way to make groups of units comparable on observed characteristics. The use of the method is, however, limited to situations where the…

Methodology · Statistics 2019-06-18 Fredrik Sävje , Michael J. Higgins , Jasjeet S. Sekhon

Personal names simultaneously differentiate individuals and categorize them in ways that are important in a given society. While the natural language processing community has thus associated personal names with sociodemographic…

Computation and Language · Computer Science 2024-07-16 Vagrant Gautam , Arjun Subramonian , Anne Lauscher , Os Keyes

A variety of statistical methods for noun compound analysis are implemented and compared. The results support two main conclusions. First, the use of conceptual association not only enables a broad coverage, but also improves the accuracy.…

cmp-lg · Computer Science 2008-02-03 Mark Lauer

Multiple-systems or capture-recapture estimation are common techniques for population size estimation, particularly in the quantitative study of human rights violations. These methods rely on multiple samples from the population, along with…

Methodology · Statistics 2018-12-27 Mauricio Sadinle

Given a set $A$ of $n$ people and a set $B$ of $m \geq n$ items, with each person having a list that ranks his/her preferred items in order of preference, we want to match every person with a unique item. A matching $M$ is called popular if…

Discrete Mathematics · Computer Science 2019-10-29 Suthee Ruangwises , Toshiya Itoh

The problem addressed concerns the determination of the average number of successive attempts of guessing a word of a certain length consisting of letters with given probabilities of occurrence. Both first- and second-order approximations…

Information Theory · Computer Science 2015-06-19 Kerstin Andersson

Names are deeply tied to human identity. They can serve as markers of individuality, cultural heritage, and personal history. However, using names as a core indicator of identity can lead to over-simplification of complex identities. When…

Computation and Language · Computer Science 2025-03-11 Siddhesh Pawar , Arnav Arora , Lucie-Aimée Kaffee , Isabelle Augenstein

Computational social scientists often harness the Web as a "societal observatory" where data about human social behavior is collected. This data enables novel investigations of psychological, anthropological and sociological research…

Computers and Society · Computer Science 2016-03-15 Fariba Karimi , Claudia Wagner , Florian Lemmerich , Mohsen Jadidi , Markus Strohmaier

Given several databases containing person-specific data held by different organizations, Privacy-Preserving Record Linkage (PPRL) aims to identify and link records that correspond to the same entity/individual across different databases…

Databases · Computer Science 2022-12-13 Dinusha Vatsalan , Dimitrios Karapiperis , Vassilios S. Verykios
‹ Prev 1 3 4 5 6 7 10 Next ›