Related papers: A Two-Parameter Weibull Framework for Diagnosing T…
Transformers have become the dominant architecture in modern machine learning, yet the theoretical understanding of their training dynamics remains limited. This paper develops a rigorous mathematical framework for analyzing gradient-based…
This paper aims to explore the inherent connection among Heisenberg groups, quantum Fourier transform and (quasiprobability) distribution functions. Distribution functions for continuous and finite quantum systems are examined first as a…
We study weight-only post-training quantization (PTQ), which quantizes the weights of a large language model (LLM) without retraining, using little or no calibration data. Weight-only PTQ is crucial for reducing the memory footprint and…
As the size of the pre-trained language model (PLM) continues to increase, numerous parameter-efficient transfer learning methods have been proposed recently to compensate for the tremendous cost of fine-tuning. Despite the impressive…
In statistical process control Weibull distribution can be used to model the time between events or failures (TBE) in a process with increasing decreasing or constant failure rates. Specifically it helps in monitoring processes where the…
Transformers have achieved extraordinary success in modern machine learning due to their excellent ability to handle sequential data, especially in next-token prediction (NTP) tasks. However, the theoretical understanding of their…
In this paper we introduce, for the first time, the Weibull-Geometric distribution which generalizes the exponential-geometric distribution proposed by Adamidis and Loukas (1998). The hazard function of the last distribution is monotone…
Despite the remarkable empirical performance of Transformers, their theoretical understanding remains elusive. Here, we consider a deep multi-head self-attention network, that is closely related to Transformers yet analytically tractable.…
While task-specific finetuning of pretrained networks has led to significant empirical advances in NLP, the large size of networks makes finetuning difficult to deploy in multi-task, memory-constrained settings. We propose diff pruning as a…
Recent advancements in semi-supervised learning have focused on a more realistic yet challenging task: addressing imbalances in labeled data while the class distribution of unlabeled data remains both unknown and potentially mismatched.…
The analysis of high dimensional survival data is challenging, primarily due to the problem of overfitting which occurs when spurious relationships are inferred from data that subsequently fail to exist in test data. Here we propose a novel…
This paper introduces a new generalization of the power generalized Weibull distribution called the generalized power generalized Weibull distribution. This distribution can also be considered as a generalization of Weibull distribution.…
We observe $n$ pairs of independent (but not necessarily i.i.d.) random variables $X_{1}=(W_{1},Y_{1}),\ldots,X_{n}=(W_{n},Y_{n})$ and tackle the problem of estimating the conditional distributions $Q_{i}^{\star}(w_{i})$ of $Y_{i}$ given…
In this paper, we develop double acceptance sampling plan and group acceptance sampling plan for an inverse Weibull distribution based on a truncated life test. We consider the median lifetime of the test units as a quality parameter and…
Predicting the outcomes of quantum measurements is a cornerstone of quantum information theory and a key resource for quantum technologies. Here, we introduce a comprehensive framework for quantifying the predictability of measurements on a…
Skip connections and normalisation layers form two standard architectural components that are ubiquitous for the training of Deep Neural Networks (DNNs), but whose precise roles are poorly understood. Recent approaches such as Deep Kernel…
Compared with previous two-stream trackers, the recent one-stream tracking pipeline, which allows earlier interaction between the template and search region, has achieved a remarkable performance gain. However, existing one-stream trackers…
In both Tweedie and geometric Tweedie models, the common power parameter $p\notin(0,1)$ works as an automatic distribution selection. It mainly separates two subclasses of semicontinuous ($1<p<2$) and positive continuous ($p\geq 2$)…
In this paper a new lifetime distribution, which is called the exponentiated Weibull-geometric (EWG) distribution, is introduced. This new distribution obtained by compounding the exponentiated Weibull and geometric distributions. The EWG…
Few-shot transfer has been revolutionized by stronger pre-trained models and improved adaptation algorithms.However, there lacks a unified, rigorous evaluation protocol that is both challenging and realistic for real-world usage. In this…