Related papers: Biases in the Experimental Annotations of Protein …
Many fundamental biological processes are regulated by protein-DNA complexes called {\it synaptosomes}, which possess multiple interaction sites. Despite the critical importance of synaptosomes, the mechanisms of their formation remain not…
Proteins are responsible for the most diverse set of functions in biology. The ability to extract information from protein sequences and to predict the effects of mutations is extremely valuable in many domains of biology and medicine.…
Active learning (AL) is a training paradigm for selecting unlabeled samples for annotation to improve model performance on a test set, which is useful when only a limited number of samples can be annotated. These algorithms often work by…
Inferences about hypotheses are ubiquitous in the cognitive sciences. Bayes factors provide one general way to compare different hypotheses by their compatibility with the observed data. Those quantifications can then also be used to choose…
The folding dynamics of small single-domain proteins is a current focus of simulations and experiments. Many of these proteins are 'two-state folders', i.e. proteins that fold rather directly from the denatured state to the native state,…
Understanding protein structure-function relationships is a key challenge in computational biology, with applications across the biotechnology and pharmaceutical industries. While it is known that protein structure directly impacts protein…
Considerable insight into the functional activity of proteins and enzymes can be obtained by studying the low-energy conformational distortions that the biopolymer can sustain. We carry out the characterization of these large scale…
Citations are essential for recognizing scientific contributions, yet citation behavior is shaped by more than just relevance or quality. We analyzed approximately 255,000 refereed astronomy articles published between 2000 and 2025 to…
Statistics comes in two main flavors: frequentist and Bayesian. For historical and technical reasons, frequentist statistics have traditionally dominated empirical data analysis, and certainly remain prevalent in empirical software…
In this article, we quantitatively study, through stochastic models, the efects of several intracellular phenomena, such as cell volume growth, cell division, gene replication as well as fuctuations of available RNA polymerases and…
External information propagates in the cell mainly through signaling cascades and transcriptional activation, allowing it to react to a wide spectrum of environmental changes. High throughput experiments identify numerous molecular…
Empirical evidence demonstrates that citations received by scholarly publications follow a pattern of preferential attachment, resulting in a power-law distribution. Such asymmetry has sparked significant debate regarding the use of…
We show that in experimental atomic force microscopy studies of the lifetime distribution of mechanically stressed folded proteins the effects of externally applied fluctuations can not be distinguished from those of internally present…
Citation numbers are extensively used for assessing the quality of scientific research. The use of raw citation counts is generally misleading, especially when applied to cross-disciplinary comparisons, since the average number of citations…
Survivorship bias is the tendency to concentrate on the positive outcomes of a selection process and overlook the results that generate negative outcomes. We observe that this bias could be present in the popular MS MARCO dataset, given…
High-quality data annotation is an essential but laborious and costly aspect of developing machine learning-based software. We explore the inherent tradeoff between annotation accuracy and cost by detecting and removing minority reports --…
We introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and…
Intracellular compartmentalization of proteins underpins their function and the metabolic processes they sustain. Various mass spectrometry-based proteomics methods (subcellular spatial proteomics) now allow high throughput subcellular…
Supervised classification heavily depends on datasets annotated by humans. However, in subjective tasks such as toxicity classification, these annotations often exhibit low agreement among raters. Annotations have commonly been aggregated…
Citation analysis is widely used in research evaluation to assess the impact of scientific papers. These analyses rest on the assumption that citation decisions by authors are accurate, representing the flow of knowledge from cited to…