中文
相关论文

相关论文: Statistical Validation of Column Matching in the D…

200 篇论文

Matching entries of correlated shuffled databases have practical applications ranging from privacy to biology. In this paper, motivated by synchronization errors in the sampling of time-indexed databases, matching of random databases under…

信息论 · 计算机科学 2022-05-10 Serhat Bakirtas , Elza Erkip

The Brazilian Mathematical Olympiads for Public Schools (OBMEP) is held every year since 2005. In the 2013 edition there were over 47,000 schools registered involving nearly 19.2 million students. The Brazilian public educational system is…

应用统计 · 统计学 2015-07-03 Alexandra M. Schmidt , Caroline P. de Moraes , Helio S. Migon

In scientific inference problems, the underlying statistical modeling assumptions have a crucial impact on the end results. There exist, however, only a few automatic means for validating these fundamental modelling assumptions. The…

统计方法学 · 统计学 2019-05-21 Andreas Svensson , Dave Zachariah , Petre Stoica , Thomas B. Schön

A new method, with an application program in Matlab code, is proposed for testing item performance models on empirical databases. This method uses data intraclass correlation statistics as expected correlations to which one compares simple…

统计方法学 · 统计学 2011-04-13 Pierre Courrieu , Muriele Brand-D'Abrescia , Ronald Peereman , Daniel Spieler , Arnaud Rey

Nowadays, data analysis in the world of Big Data is connected typically to data mining, descriptive or exploratory statistics, e.~g.\ cluster analysis, classification or regression analysis. Aside these techniques there is a huge area of…

应用统计 · 统计学 2018-10-24 Taras Lazariv , Christoph Lehmann

Sentiment Analysis is one of the most classical and primarily studied natural language processing tasks. This problem had a notable advance with the proposition of more complex and scalable machine learning models. Despite this progress,…

计算与语言 · 计算机科学 2021-12-13 Frederico Souza , João Filho

Scientific knowledge cannot be seen as a set of isolated fields, but as a highly connected network. Understanding how research areas are connected is of paramount importance for adequately allocating funding and human resources (e.g.,…

数字图书馆 · 计算机科学 2021-04-09 Francisco Galuppo Azevedo , Fabricio Murai

The fragmentation of public data in Brazil, coupled with inconsistent standards and limited interoperability, hinders effective research, evidence-based policymaking and access to data-driven insights. To address these issues, we introduce…

计算机与社会 · 计算机科学 2025-11-18 Isadora Cristina , Ramon Gonze , Jônatas Santos , Julio Reis , Mário Alvim , Bernardo Queiroz , Fabrício Benevenuto

As organizations continue to access diverse datasets, the demand for effective data integration has increased. Key tasks in this process, such as schema matching and entity resolution, are essential but often require significant effort.…

数据库 · 计算机科学 2025-11-13 Yuka Haruki , Shigeru Ishikura , Kazuya Demachi , Teruaki Hayashi

The emergence of new data sources and statistical methods is driving an update in the traditional official statistics paradigm. As an example, the Italian National Institute of Statistics (ISTAT) is undergoing a significant modernisation of…

统计方法学 · 统计学 2025-02-17 Nina Deliu , Piero Demetrio Falorsi , Stefano Falorsi , Diego Chianella , Giorgio Alleva

Quantifying the similarity between datasets has widespread applications in statistics and machine learning. The performance of a predictive model on novel datasets, referred to as generalizability, depends on how similar the training and…

统计方法学 · 统计学 2025-06-18 Marieke Stolte , Franziska Kappenberg , Jörg Rahnenführer , Andrea Bommert

The advent of high dimensional single cell data in the biomedical sciences has necessitated the development of dimensionality-reduction tools. t-SNE and UMAP are the two most frequently used approaches, allowing clear visualisation of…

A powerful approach to detecting erroneous data is to check which potentially dirty data records are incompatible with a user's domain knowledge. Previous approaches allow the user to specify domain knowledge in the form of logical…

数据库 · 计算机科学 2019-02-27 Jing Nathan Yan , Oliver Schulte , Jiannan Wang , Reynold Cheng

Now-a-days the amount of data stored in educational database increasing rapidly. These databases contain hidden information for improvement of students' performance. The performance in higher education in India is a turning point in the…

信息检索 · 计算机科学 2012-01-18 Brijesh Kumar Bhardwaj , Saurabh Pal

While opinion dynamics models have been extensively studied as stylized models, there has been growing attention to the possibility of combining these models with empirical data. This attention seems to be driven by the many social issues…

物理与社会 · 物理学 2026-02-03 Samuel Moor-Smith , Dino Carpentras

The scores obtained by students that have performed the ENEM exam, the Brazilian High School National Examination used to admit students at the Brazilian universities, is analyzed. The average high school's scores are compared between…

物理与社会 · 物理学 2016-05-04 Roberto da Silva , Luis C. Lamb , Marcia C. Barbosa

Language models are increasingly used in Brazil, but most evaluation remains English-centric. This paper presents Alvorada-Bench, a 4,515-question, text-only benchmark drawn from five Brazilian university entrance examinations. Evaluating…

计算与语言 · 计算机科学 2025-08-25 Henrique Godoy

De-anonymizing user identities by matching various forms of user data available on the internet raises privacy concerns. A fundamental understanding of the privacy leakage in such scenarios requires a careful study of conditions under which…

信息论 · 计算机科学 2021-05-21 Serhat Bakirtas , Elza Erkip

Most statistical classifiers are designed to find patterns in data where numbers fit into rows and columns, like in a spreadsheet, but many kinds of data do not conform to this structure. To uncover patterns in non-conforming data, we…

定量方法 · 定量生物学 2021-03-22 Jared Ostmeyer , Scott Christley , Lindsay Cowell

Efficient automatic protein classification is of central importance in genomic annotation. As an independent way to check the reliability of the classification, we propose a statistical approach to test if two sets of protein domain…

‹ 上一页 1 2 3 10 下一页 ›