English
Related papers

Related papers: Calculating complexity of large randomized librari…

200 papers

In many empirical studies of a large two-sided matching market (such as in a college admissions problem), the researcher performs statistical inference under the assumption that they observe a random sample from a large matching market. In…

Econometrics · Economics 2024-04-02 Jacob Schwartz , Kyungchul Song

Maximum diversity problems arise in many practical settings from facility location to social networks, and constitute an important class of NP-hard problems in combinatorial optimization. There has been a growing interest in these problems…

Optimization and Control · Mathematics 2024-01-25 Francisco Parreño , Ramón Álvarez-Valdés , Rafael Martí

Bibliographic metrics are commonly utilized for evaluation purposes within academia, often in conjunction with other metrics. These metrics vary widely across fields and change with the seniority of the scholar; consequently, the only way…

Digital Libraries · Computer Science 2021-12-17 Sen Tian , Panos Ipeirotis

The Lean mathematical library Mathlib is one of the fastest-growing libraries of formalised mathematics. We describe various strategies to manage this growth, while allowing for change and avoiding maintainer overload. This includes dealing…

Programming Languages · Computer Science 2025-10-08 Anne Baanen , Matthew Robert Ballard , Johan Commelin , Bryan Gin-ge Chen , Michael Rothgang , Damiano Testa

A density matrix describes the statistical state of a quantum system. It is a powerful formalism to represent both the quantum and classical uncertainty of quantum systems and to express different statistical operations such as measurement,…

Machine Learning · Computer Science 2024-05-01 Fabio A. González , Alejandro Gallego , Santiago Toledo-Cortés , Vladimir Vargas-Calderón

We present Cryptomite, a Python library of randomness extractor implementations. The library offers a range of two-source, seeded and deterministic randomness extractors, together with parameter calculation modules, making it easy to use…

Cryptography and Security · Computer Science 2025-11-27 Cameron Foreman , Richie Yeung , Alec Edgington , Florian J. Curchod

To assist in the development of machine learning methods for automated classification of spectroscopic data, we have generated a universal synthetic dataset that can be used for model validation. This dataset contains artificial spectra…

Machine Learning · Computer Science 2022-06-15 Jan Schuetzke , Nathan J. Szymanski , Markus Reischl

We wish to estimate the total number of classes in a population based on sample counts, especially in the presence of high latent diversity. Drawing on probability theory that characterizes distributions on the integers by ratios of…

Methodology · Statistics 2014-12-10 A. Willis , J. Bunge

Binary code is pervasive, and binary analysis is a key task in reverse engineering, malware classification, and vulnerability discovery. Unfortunately, while there exist large corpora of malicious binaries, obtaining high-quality corpora of…

Cryptography and Security · Computer Science 2024-11-05 Chang Liu , Rebecca Saul , Yihao Sun , Edward Raff , Maya Fuchs , Townsend Southard Pantano , James Holt , Kristopher Micinski

Atomic-level simulations are widely used to study biomolecules and their dynamics. A common goal in such studies is to compare simulations of a molecular system under several conditions -- for example, with various mutations or bound…

Biomolecules · Quantitative Biology 2025-01-07 Martin Vögele , Neil J. Thomson , Sang T. Truong , Jasper McAvity , Ulrich Zachariae , Ron O. Dror

Protein activity is a significant characteristic for recombinant proteins which can be used as biocatalysts. High activity of proteins reduces the cost of biocatalysts. A model that can predict protein activity from amino acid sequence is…

Quantitative Methods · Quantitative Biology 2018-07-23 X. Han , X. Wang , K. Zhou

Large matrices arise in many machine learning and data analysis applications, including as representations of datasets, graphs, model weights, and first and second-order derivatives. Randomized Numerical Linear Algebra (RandNLA) is an area…

Machine Learning · Computer Science 2024-06-21 Michał Dereziński , Michael W. Mahoney

A statistical estimation algorithm of the weight distribution of a linear code is shown, based on using its generator matrix as a compression function on random bit strings.

Information Theory · Computer Science 2018-06-07 Alessandro Tomasi , Alessio Meneghetti

Computational notebooks, such as Jupyter notebooks, are interactive computing environments that are ubiquitous among data scientists to perform data wrangling and analytic tasks. To measure the performance of AI pair programmers that…

Most protocols for the high-throughput directed evolution of enzymes rely on random encapsulation to link phenotype and genotype. In order to optimize these approaches, or compare one to another, one needs a measure of their performance at…

Populations and Evolution · Quantitative Biology 2018-11-14 Adèle Dramé-Maigné , Anton Zadorin , Iaroslava Golovkova , Yannick Rondelez

This article describes the R package moodlequizR, which allows the user to easily create fully randomized quizzes and exams for Moodle, or indeed any online assessment platform that uses XML files for importing questions. In such a quiz the…

Applications · Statistics 2024-11-13 Wolfgang Rolke

The emergence of a predominant phenotype within a cell population is often triggered by a rare accumulation of DNA mutations in a single cell. For example, tumors may be initiated by a single cell in which multiple mutations cooperate to…

Tissues and Organs · Quantitative Biology 2018-11-21 Philip Greulich , Benjamin D. Simons

The outcome of a functional genomics pipeline is usually a partial list of genomic features, ranked by their relevance in modelling biological phenotype in terms of a classification or regression model. Due to resampling protocols or just…

Machine Learning · Statistics 2012-08-20 Giuseppe Jurman , Samantha Riccadonna , Roberto Visintainer , Cesare Furlanello

The number partitioning problem consists of partitioning a sequence of positive numbers ${a_1,a_2,..., a_N}$ into two disjoint sets, ${\cal A}$ and ${\cal B}$, such that the absolute value of the difference of the sums of $a_j$ over the two…

Statistical Mechanics · Physics 2009-10-31 F. F. Ferreira , J. F. Fontanari

Science advances not only through the accumulation of facts but also through the evolution of tools. Crucially, tools are rarely used in isolation. They form tool portfolios, combinations shaped by a discipline's workflows and analytical…

Digital Libraries · Computer Science 2026-04-28 Zhouming Wu , Dakota Murray