中文
相关论文

相关论文: Learning the language of QCD jets with transformer…

200 篇论文

We compare the performance of a convolutional neural network (CNN) trained on jet images with dense neural networks (DNNs) trained on n-subjettiness variables to study the distinguishing power of these two separate techniques applied to top…

高能物理 - 唯象学 · 物理学 2019-09-25 Liam Moore , Karl Nordström , Sreedevi Varma , Malcolm Fairbairn

Jet quenching and more generally physics at high transverse momentum P_T scales is a cornerstone of the heavy-ion physics program at the LHC. In this work, the current understanding of jet quenching in terms of a QCD shower evolution being…

高能物理 - 唯象学 · 物理学 2015-06-17 Thorsten Renk

Scientists often use observational time series data to study complex natural processes, but regression analyses often assume simplistic dynamics. Recent advances in deep learning have yielded startling improvements to the performance of…

机器学习 · 计算机科学 2023-04-21 Cory Shain , William Schuler

Natural Language Inference is an important task for Natural Language Understanding. It is concerned with classifying the logical relation between two sentences. In this paper, we propose several text generative neural networks for…

人工智能 · 计算机科学 2017-03-28 Janez Starc , Dunja Mladenić

Recent artificial neural networks that process natural language achieve unprecedented performance in tasks requiring sentence-level understanding. As such, they could be interesting models of the integration of linguistic information in the…

计算与语言 · 计算机科学 2023-02-17 Sophie Arana , Jacques Pesnot Lerousseau , Peter Hagoort

The remarkable capability of over-parameterised neural networks to generalise effectively has been explained by invoking a ``simplicity bias'': neural networks prevent overfitting by initially learning simple classifiers before progressing…

计算与语言 · 计算机科学 2025-10-02 Riccardo Rende , Federica Gerace , Alessandro Laio , Sebastian Goldt

Quantile regression is increasingly encountered in modern big data applications due to its robustness and flexibility. We consider the scenario of learning the conditional quantiles of a specific target population when the available data…

统计理论 · 数学 2024-02-27 Jun Jin , Jun Yan , Robert H. Aseltine , Kun Chen

Understanding the training dynamics of transformers is important to explain the impressive capabilities behind large language models. In this work, we study the dynamics of training a shallow transformer on a task of recognizing…

机器学习 · 计算机科学 2024-10-15 Hongru Yang , Bhavya Kailkhura , Zhangyang Wang , Yingbin Liang

Large language models based on transformer architectures have become integral to state-of-the-art natural language processing applications. However, their training remains computationally expensive and exhibits instabilities, some of which…

数值分析 · 数学 2025-03-14 Stanislav Budzinskiy , Wenyi Fang , Longbin Zeng , Philipp Petersen

The Large Hadron Collider (LHC) is the first machine that provides high enough energy to produce large numbers of boosted top quarks. The decay products of these top quarks are confined to a cone in the top quark flight direction and can be…

高能物理 - 实验 · 物理学 2015-09-14 Sebastian Schätzel

A good feature representation is a determinant factor to achieve high performance for many machine learning algorithms in terms of classification. This is especially true for techniques that do not build complex internal representations of…

神经与进化计算 · 计算机科学 2019-08-22 Noëlie Cherrier , Jean-Philippe Poli , Maxime Defurne , Franck Sabatié

We demonstrate the emergence of scaling laws in the benchmark top versus QCD jet classification problem in collider physics. Six distinct physically-motivated classifiers exhibit power-law scaling of the binary cross-entropy test loss as a…

高能物理 - 唯象学 · 物理学 2023-12-06 Joshua Batson , Yonatan Kahn

Natural Language Generation (NLG) models are prone to generating repetitive utterances. In this work, we study the repetition problem for encoder-decoder models, using both recurrent neural network (RNN) and transformer architectures. To…

计算与语言 · 计算机科学 2020-04-10 Shaojie Jiang , Thomas Wolf , Christof Monz , Maarten de Rijke

Recent literature on deep neural networks for tagging of highly energetic jets resulting from top quark decays has focused on image based techniques or multivariate approaches using high-level jet substructure variables. Here, a sequential…

高能物理 - 实验 · 物理学 2017-08-10 Jannicke Pearkes , Wojciech Fedorko , Alison Lister , Colin Gay

We describe a novel end-to-end approach using Machine Learning to reconstruct the power spectrum of cosmological density perturbations at high redshift from observed quasar spectra. State-of-the-art cosmological simulations of structure…

宇宙学与河外天体物理 · 物理学 2021-07-21 Maria Han Veiga , Xi Meng , Oleg Y. Gnedin , Nickolay Y. Gnedin , Xun Huan

We present a transformer architecture-based foundation model for tasks at high-energy particle colliders such as the Large Hadron Collider. We train the model to classify jets using a self-supervised strategy inspired by the Joint Embedding…

机器学习 · 计算机科学 2025-02-07 Jai Bardhan , Radhikesh Agrawal , Abhiram Tilak , Cyrin Neeraj , Subhadip Mitra

The search for new physics at high energy accelerators has been at the crossroads with very little hint of signals suggesting otherwise. The challenges at a hadronic machine such as the LHC is compounded by the fact that final states are…

高能物理 - 唯象学 · 物理学 2024-06-12 Aruna Kumar Nayak , Santosh Kumar Rai , Tousik Samui

Large Language Models (LLMs) have the capacity to store and recall facts. Through experimentation with open-source models, we observe that this ability to retrieve facts can be easily manipulated by changing contexts, even without altering…

计算与语言 · 计算机科学 2024-12-02 Yibo Jiang , Goutham Rajendran , Pradeep Ravikumar , Bryon Aragam

This research note combines two methods that have recently improved the state of the art in language modeling: Transformers and dynamic evaluation. Transformers use stacked layers of self-attention that allow them to capture long range…

机器学习 · 计算机科学 2019-04-18 Ben Krause , Emmanuel Kahembwe , Iain Murray , Steve Renals

Large, pre-trained neural networks consisting of self-attention layers (transformers) have recently achieved state-of-the-art results on several speech emotion recognition (SER) datasets. These models are typically pre-trained in…