Related papers: Analyzing LC-MS/MS data by spectral count and ion …
We compare two deletion-based methods for dealing with the problem of missing observations in linear regression analysis. One is the complete-case analysis (CC, or listwise deletion) that discards all incomplete observations and only uses…
Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for efficient data selection methods that reduce annotation costs…
Elemental abundance diagnostics in the solar corona are crucial for understanding energy transport, plasma heating, and magnetic activity. Most earlier imaging-spectroscopic studies have relied on slit-based spectrometers, which offer high…
Quantifying the similarity of two or more datasets has widespread applications in statistics and machine learning. The method choice is, however, difficult due to the abundance of proposed methods and the lack of neutral comparison studies,…
Missing data is a common issue in many biomedical studies. Under a paired design, some subjects may have missing values in either one or both of the conditions due to loss of follow-up, insufficient biological samples, etc. Such partially…
Large-scale hypothesis testing has become a ubiquitous problem in high-dimensional statistical inference, with broad applications in various scienfitic disciplines. One relevant application is constituted by imaging mass spectrometry (IMS)…
We consider equilibrium folding transitions in lattice protein models with and without side chains. A dimensionless measure, $Omega_{c}$, is introduced to quantitatively assess the degree of cooperativity in lattice models and in real…
The coupling of an electron monochromator (EM) to a mass spectrometer (MS) has created a new analytical technique, EM-MS, for the investigation of electrophilic compounds. This method provides a powerful tool for molecular identification of…
Machine learning (ML) methods have proved to be a very successful tool in physical sciences, especially when applied to experimental data analysis. Artificial intelligence is particularly good at recognizing patterns in high dimensional…
The following article presents a multi-length-scale characterization approach for investigating doping chemistry and spatial distributions within semiconductors, as demonstrated using a state-of-the-art CMOS image sensor. With an intricate…
Numerous types of social biases have been identified in pre-trained language models (PLMs), and various intrinsic bias evaluation measures have been proposed for quantifying those social biases. Prior works have relied on human annotated…
Data in biology is redundant, noisy, and sparse. How does the type and scale of available data impact model performance? In this work, we specifically investigate how protein language models (pLMs) scale with increasing pretraining data. We…
Robust model-fitting to spectroscopic transitions is a requirement across many fields of science. The corrected Akaike and Bayesian information criteria (AICc and BIC) are most frequently used to select the optimal number of fitting…
Motivation: Tumor classification using Imaging Mass Spectrometry (IMS) data has a high potential for future applications in pathology. Due to the complexity and size of the data, automated feature extraction and classification steps are…
Due to its specificity, fluorescence microscopy (FM) has become a quintessential imaging tool in cell biology. However, photobleaching, phototoxicity, and related artifacts continue to limit FM's utility. Recently, it has been shown that…
This chapter reviews recent developments in the use of mixed-species ion chains in quantum information science, frequency metrology and spectroscopy. A growing number of experiments have demonstrated new methods in this area, opening up new…
In light of newly developed standardization methods, we evaluate, via simulation study, how propensity score weighting and standardization -based approaches compare for obtaining estimates of the marginal odds ratio and the marginal hazard…
We report the preliminary results of a meta-analysis conducted to examine possible biases in the uncertainty values published in papers by ATLAS and CMS experiments. We have performed this analysis using two independent techniques; a…
In SPECT, list-mode (LM) format allows storing data at higher precision compared to binned data. There is significant interest in investigating whether this higher precision translates to improved performance on clinical tasks. Towards this…
While data science is battling to extract information from the enormous explosion of data, many estimators and algorithms are being developed for better prediction. Researchers and data scientists often introduce new methods and evaluate…