Related papers: Multi-scale cross-attention transformer encoder fo…
The physics potential of massive liquid argon TPCs in the low-energy regime is still to be fully reaped because few-hits events encode information that can hardly be exploited by conventional classification algorithms. Machine learning (ML)…
At the extreme energies of the Large Hadron Collider, massive particles can be produced at such high velocities that their hadronic decays are collimated and the resulting jets overlap. Deducing whether the substructure of an observed jet…
We propose AttentionMixer, a unified deep learning framework for multimodal detection of brain edema that combines structural head CT (HCT) with routine clinical metadata. While HCT provides rich spatial information, clinical variables such…
A search is presented for new particles in an extension to the Standard Model that includes a heavy Higgs boson H0, an intermediate charged Higgs boson pair H+-, and a light Higgs boson h0. The analysis searches for events involving the…
Survival prediction is crucial for cancer patients as it provides early prognostic information for treatment planning. Recently, deep survival models based on deep learning and medical images have shown promising performance for survival…
We present a novel technique for the analysis of proton-proton collision events from the ATLAS and CMS experiments at the Large Hadron Collider. For a given final state and choice of kinematic variables, we build a graph network in which…
Identifying and reconstructing hadronic $\tau$ decays ($\tau_{\textrm{h}}$) is an important task at current and future high-energy physics experiments, as $\tau_{\textrm{h}}$ represent an important tool to analyze the production of Higgs…
Human state recognition is a critical topic with pervasive and important applications in human-machine systems. Multi-modal fusion, the combination of metrics from multiple data sources, has been shown as a sound method for improving the…
The shape of the Higgs potential is modified by the presence of additional scalar fields, as predicted in many Beyond-Standard-Model (BSM) scenarios. In such cases, deviations in the Higgs self-interactions, in particular the trilinear…
An inclusive search for the standard model Higgs boson ($\mathrm{H}$) produced with large transverse momentum ($p_\mathrm{T}$) and decaying to a bottom quark-antiquark pair ($\mathrm{b}\overline{\mathrm{b}}$) is performed using a data set…
Non-hierarchical sparse attention Transformer-based models, such as Longformer and Big Bird, are popular approaches to working with long documents. There are clear benefits to these approaches compared to the original Transformer in terms…
This study introduces an innovative approach to analyzing unlabeled data in high-energy physics (HEP) through the application of self-supervised learning (SSL). Faced with the increasing computational cost of producing high-quality labeled…
Transformers have become prevalent in computer vision due to their performance and flexibility in modelling complex operations. Of particular significance is the 'cross-attention' operation, which allows a vector representation (e.g. of an…
The transformer structure employed in large language models (LLMs), as a specialized category of deep neural networks (DNNs) featuring attention mechanisms, stands out for their ability to identify and highlight the most relevant aspects of…
Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. Motivated by the efficiency gains of hybrid models and the broad availability of pretrained large…
Extending the Standard Model (SM) by a $U(1)_{L_\mu-L_\tau}$ group gives potentially significant new contributions to $g_\mu-2$, allows the construction of realistic neutrino mass matrices, incorporates lepton universality violation, and…
The sensitivity of the LHC experiments to the Standard Model Higgs using $H\to\tau\tau\to l^+l^-\sla{p_t}$ associated with one high $P_T$ jet in the mass range $110 <M_H<150 \gev$/c$^2$ is investigated. A cut and Neural Network based event…
Large Language Models (LLMs) based on autoregressive, decoder-only Transformers generate text one token at a time, where a token represents a discrete unit of text. As each newly produced token is appended to the partial output sequence,…
Speculative decoding (SD) is a widely adopted approach for accelerating inference in large language models (LLMs), particularly when the draft and target models are well aligned. However, state-of-the-art SD methods typically rely on…
We examine the discovery potential for double Higgs production at the high luminosity LHC in the final state with two $b$-tagged jets, two leptons and missing transverse momentum. Although this dilepton final state has been considered a…