English
Related papers

Related papers: Robust estimation of isoform expression with RNA-S…

200 papers

We focus on the estimation of the intensity of a Poisson process in the presence of a uniform noise. We propose a kernel-based procedure fully calibrated in theory and practice. We show that our adaptive estimator is optimal from the oracle…

Methodology · Statistics 2022-06-29 Anna Bonnet , Claire Lacour , Franck Picard , Vincent Rivoirard

Recent advances in molecular biology allow the quantification of the transcriptome and scoring transcripts as differentially or equally expressed between two biological conditions. Although these two tasks are closely linked, the available…

Methodology · Statistics 2017-02-08 Panagiotis Papastamoulis , Magnus Rattray

Estimating the difficulty of input questions as perceived by large language models (LLMs) is essential for accurate performance evaluation and adaptive inference. Existing methods typically rely on repeated response sampling, auxiliary…

Computation and Language · Computer Science 2025-09-17 Yubo Zhu , Dongrui Liu , Zecheng Lin , Wei Tong , Sheng Zhong , Jing Shao

Existing sequential generalized estimating equation methodology for longitudinal and group-correlated data focuses on narrow hypotheses concerning treatment efficacy and often makes modeling assumptions that impede the desirable robustness…

Methodology · Statistics 2026-03-16 Nathan T. Provost , Abdus S. Wahed

In this paper, we propose an outlier-robust regularized kernel-based method for linear system identification. The unknown impulse response is modeled as a zero-mean Gaussian process whose covariance (kernel) is given by the recently…

Machine Learning · Statistics 2013-12-25 Giulio Bottegal , Aleksandr Y. Aravkin , Hakan Hjalmarsson , Gianluigi Pillonetto

The accurate quantification of gene expression levels is crucial for transcriptome study. Microarray platforms are commonly used for simultaneously interrogating thousands of genes in the past decade, and recently RNA-Seq has emerged as a…

Applications · Statistics 2014-08-01 Zhaonan Sun , Thomas Kuczek , Yu Zhu

Model mis-specification (e.g. the presence of outliers) is commonly encountered in astronomical analyses, often requiring the use of ad hoc algorithms which are sensitive to arbitrary thresholds (e.g. sigma-clipping). For any given dataset,…

Instrumentation and Methods for Astrophysics · Physics 2025-09-03 William Martin , Daniel J. Mortlock

The widespread adoption of large language models (LLMs) necessitates reliable methods to detect LLM-generated text. We introduce SimMark, a robust sentence-level watermarking algorithm that makes LLMs' outputs traceable without requiring…

Computation and Language · Computer Science 2025-09-12 Amirhossein Dabiriaghdam , Lele Wang

In genomics studies, the investigation of the gene relationship often brings important biological insights. Currently, the large heterogeneous datasets impose new challenges for statisticians because gene relationships are often local. They…

Methodology · Statistics 2022-03-07 Jinjin Tian , Jing Lei , Kathryn Roeder

Unraveling the co-expression of genes across studies enhances the understanding of cellular processes. Inferring gene co-expression networks from transcriptome data presents many challenges, including spurious gene correlations, sample…

Machine Learning · Statistics 2024-10-01 Teodora Pandeva , Martijs Jonker , Leendert Hamoen , Joris Mooij , Patrick Forré

We develop an efficient posterior sampling scheme for the Poisson INGARCH models. The proposed method is based on the approximation of the posterior density that exploits the Poisson limit of the negative binomial distribution. It allows us…

Methodology · Statistics 2026-03-10 Yixuan Fan , Zhengwei Liu , Fukang Zhu

In this report a systematic approach is used to determine the approximate genetic network and robust dependencies underlying differentiation. The data considered is in the form of a binary matrix and represent the expression of the nine…

Molecular Networks · Quantitative Biology 2007-05-23 Radhakrishnan Nagarajan , Jane E. Aubin , Charlotte A. Peterson

We propose a probabilistic model for interpreting gene expression levels that are observed through single-cell RNA sequencing. In the model, each cell has a low-dimensional latent representation. Additional latent variables account for…

Machine Learning · Computer Science 2017-10-18 Romain Lopez , Jeffrey Regier , Michael Cole , Michael Jordan , Nir Yosef

Sampling is a promising bottom-up method for exposing what generative models have learned about language, but it remains unclear how to generate representative samples from popular masked language models (MLMs) like BERT. The MLM objective…

Computation and Language · Computer Science 2022-03-21 Takateru Yamakoshi , Thomas L. Griffiths , Robert D. Hawkins

Clustering with variable selection is a challenging yet critical task for modern small-n-large-p data. Existing methods based on sparse Gaussian mixture models or sparse K-means provide solutions to continuous data. With the prevalence of…

Machine Learning · Statistics 2020-04-28 Tanbin Rahman , Yujia Li , Tianzhou Ma , Lu Tang , George Tseng

A novel sequential inferential method for Bayesian dynamic generalised linear models is presented, addressing both univariate and multivariate $k$-parametric exponential families. It efficiently handles diverse responses, including…

Methodology · Statistics 2025-01-15 Mariane Branco Alves , Helio S. Migon , Silvaneo V. Santos , Raíra Marotta

This study investigates the structured generation capabilities of large language models (LLMs), focusing on producing valid JSON outputs against a given schema. Despite the widespread use of JSON in integrating language models with…

Computation and Language · Computer Science 2025-03-07 Yaxi Lu , Haolun Li , Xin Cong , Zhong Zhang , Yesai Wu , Yankai Lin , Zhiyuan Liu , Fangming Liu , Maosong Sun

In this paper, we consider the situation in which the observations follow an isotonic generalized partly linear model. Under this model, the mean of the responses is modelled, through a link function, linearly on some covariates and…

Statistics Theory · Mathematics 2018-11-30 Graciela Boente , Daniela Rodriguez , Pablo Vena

We study the problem of identifying change points in high-dimensional generalized linear models, and propose an approach based on sample-weighted empirical risk minimization. Our method, Weighted ERM, encodes priors on the change points via…

Methodology · Statistics 2026-04-14 Gabriel Arpino , Ramji Venkataramanan

Despite their outstanding performance, large language models (LLMs) suffer notorious flaws related to their preference for simple, surface-level textual relations over full semantic complexity of the problem. This proposal investigates a…

Computation and Language · Computer Science 2022-06-20 Michal Štefánik