中文
相关论文

相关论文: Learning the language of QCD jets with transformer…

200 篇论文

Transformer-based language models are effective but complex, and understanding their inner workings and reasoning mechanisms is a significant challenge. Previous research has primarily explored how these models handle simple tasks like name…

计算与语言 · 计算机科学 2025-05-20 Zeyuan Allen-Zhu , Yuanzhi Li

We study high-pt jets from QCD and from highly-boosted massive particles such as tops, W, Z and Higgs, and argue that infrared-safe observables can help reduce QCD backgrounds. Jets from QCD are characterized by different patterns of energy…

高能物理 - 唯象学 · 物理学 2014-11-18 Leandro G. Almeida , Seung J. Lee , Gilad Perez , George Sterman , Ilmo Sung , Joseph Virzi

Natural Language Processing (NLP) relies heavily on training data. Transformers, as they have gotten bigger, have required massive amounts of training data. To satisfy this requirement, text augmentation should be looked at as a way to…

计算与语言 · 计算机科学 2022-11-17 Matthew Ciolino , David Noever , Josh Kalin

We introduce a potentially powerful new method of searching for new physics at the LHC, using autoencoders and unsupervised deep learning. The key idea of the autoencoder is that it learns to map "normal" events back to themselves, but…

高能物理 - 唯象学 · 物理学 2020-04-22 Marco Farina , Yuichiro Nakai , David Shih

Diffusion models have achieved great success in modeling continuous data modalities such as images, audio, and video, but have seen limited use in discrete domains such as language. Recent attempts to adapt diffusion to language have…

计算与语言 · 计算机科学 2023-11-08 Justin Lovelace , Varsha Kishore , Chao Wan , Eliot Shekhtman , Kilian Q. Weinberger

In natural language processing, current methods for understanding Transformers are successful at identifying intermediate predictions during a model's inference. However, these approaches function as limited diagnostic checkpoints, lacking…

机器学习 · 计算机科学 2025-12-18 Aditya Gupta , Kirandeep Kaur , Vinayak Gupta , Chirag Shah

Transformers can generate predictions in two approaches: 1. auto-regressively by conditioning each sequence element on the previous ones, or 2. directly produce an output sequences in parallel. While research has mostly explored upon this…

计算机视觉与模式识别 · 计算机科学 2021-08-18 Andrea Alfieri , Yancong Lin , Jan C. van Gemert

Jet classification in high-energy particle physics is important for understanding fundamental interactions and probing phenomena beyond the Standard Model. Jets originate from the fragmentation and hadronization of quarks and gluons, and…

数据分析、统计与概率 · 物理学 2025-08-15 Juvenal Bassa , Vidya Manian , Sudhir Malik , Arghya Chattopadhyay

Despite the empirical success of prompt tuning in adapting pretrained language models to new tasks, theoretical analyses of its capabilities remain limited. Existing theoretical work primarily addresses universal approximation properties,…

机器学习 · 计算机科学 2025-09-03 Maxime Meyer , Mario Michelessa , Caroline Chaux , Vincent Y. F. Tan

The suppression and modification of high-energy objects, like jets, in heavy-ion collisions provide an important window to access the degrees of freedom of the quark-gluon plasma on different length scales. Despite increasingly precise and…

核理论 · 物理学 2021-01-01 Jasmine Brewer

A language model (LM) is a mapping from a linguistic context to an output token. However, much remains to be known about this mapping, including how its geometric properties relate to its function. We take a high-level geometric approach to…

计算与语言 · 计算机科学 2025-05-01 Emily Cheng , Diego Doimo , Corentin Kervadec , Iuri Macocco , Jade Yu , Alessandro Laio , Marco Baroni

Words in natural language follow a Zipfian distribution whereby some words are frequent but most are rare. Learning representations for words in the "long tail" of this distribution requires enormous amounts of data. Representations of rare…

With the rapid development of AI technology in recent years, there have been many studies with deep learning models in soft sensing area. However, the models have become more complex, yet, the data sets remain limited: researchers are…

机器学习 · 计算机科学 2022-01-25 Chao Zhang , Jaswanth Yella , Yu Huang , Xiaoye Qian , Sergei Petrov , Andrey Rzhetsky , Sthitie Bom

Data augmentation methods for Natural Language Processing tasks are explored in recent years, however they are limited and it is hard to capture the diversity on sentence level. Besides, it is not always possible to perform data…

计算与语言 · 计算机科学 2022-05-20 M. Şafak Bilici , Mehmet Fatih Amasyali

Wavelets have emerged as a cutting edge technology in a number of fields. Concrete results of their application in Image and Signal processing suggest that wavelets can be effectively applied to Natural Language Processing (NLP) tasks that…

计算与语言 · 计算机科学 2025-08-04 Rana Salama , Abdou Youssef , Mona Diab

Although deep convolutional networks have achieved improved performance in many natural language tasks, they have been treated as black boxes because they are difficult to interpret. Especially, little is known about how they represent…

计算与语言 · 计算机科学 2019-03-01 Seil Na , Yo Joong Choe , Dong-Hyun Lee , Gunhee Kim

We show that deep learning models, and especially architectures like the Transformer, originally intended for natural language, can be trained on randomly generated datasets to predict to very high accuracy both the qualitative and…

机器学习 · 计算机科学 2021-12-08 François Charton , Amaury Hayat , Sean T. McQuade , Nathaniel J. Merrill , Benedetto Piccoli

Hard probes are a cornerstone in the ongoing program to determine the properties of hot and dense QCD matter as created in ultrarelativistic heavy ion collisions. LHC measurements have so far resulted in a wealth of high P_T data, opening…

高能物理 - 唯象学 · 物理学 2015-06-11 Thorsten Renk

Understanding the properties of the quark-gluon plasma (QGP) that is produced in ultra-relativistic nucleus-nucleus collisions has been one of the top priorities of the heavy ion program at the LHC. Energetic jets are produced and…

高能物理 - 唯象学 · 物理学 2015-06-23 Yang-Ting Chien

We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a $2$ trillion token database, our Retrieval-Enhanced Transformer (RETRO)…