English
Related papers

Related papers: Decomposition of Small Transformer Models

200 papers

The aim of this paper is to discuss potential advances in PET kinetic models and direct reconstruction of kinetic parameters. As a prominent example we focus on a typical task in perfusion imaging and derive a system of…

Optimization and Control · Mathematics 2014-11-20 Louise Reips , Martin Burger , Ralf Engbers

Time-domain Transformer neural networks have proven their superiority in speech separation tasks. However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion. In this…

Sound · Computer Science 2022-07-01 Jian Luo , Jianzong Wang , Ning Cheng , Edward Xiao , Xulong Zhang , Jing Xiao

Domain decomposition is a technique used to reduce memory overhead on large neutron transport problems. Currently, the optimal load-balanced processor allocation for these domains is typically determined through small-scale simulations of…

Computational Physics · Physics 2025-08-18 Alexander Mote , Todd Palmer , Lizhong Chen

Large Eddy Simulation is based on decomposition of turbulent flow structures to large energy containing scales and small subgrid scales. The present study captures all flow low energy modes of a sample shear layer using the proper…

Fluid Dynamics · Physics 2018-05-30 Hossein Rahmani , Hamid Kalaei , Ghasem Akbari , Nader Montazerin

Feature transformation plays a critical role in enhancing machine learning model performance by optimizing data representations. Recent state-of-the-art approaches address this task as a continuous embedding optimization problem, converting…

Machine Learning · Computer Science 2025-08-29 Yang Gao , Dongjie Wang , Scott Piersall , Ye Zhang , Liqiang Wang

In this paper, we propose a novel learning-aided sphere decoding (SD) scheme for large multiple-input--multiple-output systems, namely, deep path prediction-based sphere decoding (DPP-SD). In this scheme, we employ a neural network (NN) to…

Information Theory · Computer Science 2020-01-03 Doyeon Weon , Kyungchun Lee

Deep learning models have achieved remarkable success in different areas of machine learning over the past decade; however, the size and complexity of these models make them difficult to understand. In an effort to make them more…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Vikram V. Ramaswamy , Sunnie S. Y. Kim , Nicole Meister , Ruth Fong , Olga Russakovsky

Prior knowledge about the imaging physics provides a mechanistic forward operator that plays an important role in image reconstruction, although myriad sources of possible errors in the operator could negatively impact the reconstruction…

Image and Video Processing · Electrical Eng. & Systems 2022-11-04 Maryam Toloubidokhti , Nilesh Kumar , Zhiyuan Li , Prashnna K. Gyawali , Brian Zenger , Wilson W. Good , Rob S. MacLeod , Linwei Wang

The linear spline growth model (LSGM), which approximates complex patterns using at least two linear segments, is a popular tool for examining nonlinear change patterns. Among such models, the linear-linear piecewise change pattern is the…

Methodology · Statistics 2022-05-10 Jin Liu , Robert A. Perera , Le Kang , Robert M. Kirkpatrick , Roy T. Sabo

This paper concerns the data-driven sensor deployment problem in large spatiotemporal fields. Traditionally, sensor deployment strategies have been heavily dependent on model-based planning approaches. However, model-based approaches do not…

Signal Processing · Electrical Eng. & Systems 2022-01-04 Jiahong Chen

GPT is an auto-regressive Transformer-based pre-trained language model which has attracted a lot of attention in the natural language processing (NLP) domain due to its state-of-the-art performance in several downstream tasks. The success…

Computation and Language · Computer Science 2021-10-18 Ali Edalati , Marzieh Tahaei , Ahmad Rashid , Vahid Partovi Nia , James J. Clark , Mehdi Rezagholizadeh

Mechanistic interpretability seeks to reverse engineer a trained neural network by identifying the minimal subset of internal components. We perform a mechanistic interpretability analysis of the Particle Transformer architecture, trained…

High Energy Physics - Phenomenology · Physics 2026-05-12 Saurabh Rai , Sanmay Ganguly

The Gottesman-Kitaev-Preskill (GKP) error correcting code encodes a finite dimensional logical space in one or more bosonic modes, and has recently been demonstrated in trapped ions and superconducting microwave cavities. In this work we…

Quantum Physics · Physics 2024-03-05 Mackenzie H. Shaw , Andrew C. Doherty , Arne L. Grimsmo

One of the unspoken challenges of tractography is choosing the right parameters for a given dataset or bundle. In order to tackle this challenge, we explore the multi-dimensional parameter space of tractography using streamline-specific…

Image and Video Processing · Electrical Eng. & Systems 2024-08-12 Ruben Vink , Anna Vilanova , Maxime Chamberland

Mechanistic interpretability has transformed the analysis of transformer circuits by decomposing model behavior into competing algorithms, identifying phase transitions during training, and deriving closed-form predictions for when and why…

Machine Learning · Computer Science 2026-03-19 Alma Lago

Gaussian processes (GPs) are a powerful tool for probabilistic inference over functions. They have been applied to both regression and non-linear dimensionality reduction, and offer desirable properties such as uncertainty estimates,…

Machine Learning · Statistics 2014-10-01 Yarin Gal , Mark van der Wilk , Carl E. Rasmussen

Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize two representational hypotheses: (1) a representation in the…

Recent advances in fine-grained representation learning leverage local-to-global (emergent) relationships for achieving state-of-the-art results. The relational representations relied upon by such methods, however, are abstract. We aim to…

Computer Vision and Pattern Recognition · Computer Science 2023-10-25 Abhra Chaudhuri , Massimiliano Mancini , Zeynep Akata , Anjan Dutta

Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components. Without a formal framework, however, mechanistic explanations cannot be…

Machine Learning · Computer Science 2026-05-12 Ward Gauderis , Thomas Dooms , Steven T. Holmer , Kola Ayonrinde , Geraint A. Wiggins

For an explanation of a deep learning model to be effective, it must provide both insight into a model and suggest a corresponding action in order to achieve some objective. Too often, the litany of proposed explainable deep learning…

Machine Learning · Computer Science 2020-10-09 Laura Rieger , Chandan Singh , W. James Murdoch , Bin Yu