中文
相关论文

相关论文: Decomposition of Small Transformer Models

200 篇论文

The aim of this paper is to discuss potential advances in PET kinetic models and direct reconstruction of kinetic parameters. As a prominent example we focus on a typical task in perfusion imaging and derive a system of…

最优化与控制 · 数学 2014-11-20 Louise Reips , Martin Burger , Ralf Engbers

Time-domain Transformer neural networks have proven their superiority in speech separation tasks. However, these models usually have a large number of network parameters, thus often encountering the problem of GPU memory explosion. In this…

声音 · 计算机科学 2022-07-01 Jian Luo , Jianzong Wang , Ning Cheng , Edward Xiao , Xulong Zhang , Jing Xiao

Domain decomposition is a technique used to reduce memory overhead on large neutron transport problems. Currently, the optimal load-balanced processor allocation for these domains is typically determined through small-scale simulations of…

计算物理 · 物理学 2025-08-18 Alexander Mote , Todd Palmer , Lizhong Chen

Large Eddy Simulation is based on decomposition of turbulent flow structures to large energy containing scales and small subgrid scales. The present study captures all flow low energy modes of a sample shear layer using the proper…

流体动力学 · 物理学 2018-05-30 Hossein Rahmani , Hamid Kalaei , Ghasem Akbari , Nader Montazerin

Feature transformation plays a critical role in enhancing machine learning model performance by optimizing data representations. Recent state-of-the-art approaches address this task as a continuous embedding optimization problem, converting…

机器学习 · 计算机科学 2025-08-29 Yang Gao , Dongjie Wang , Scott Piersall , Ye Zhang , Liqiang Wang

In this paper, we propose a novel learning-aided sphere decoding (SD) scheme for large multiple-input--multiple-output systems, namely, deep path prediction-based sphere decoding (DPP-SD). In this scheme, we employ a neural network (NN) to…

信息论 · 计算机科学 2020-01-03 Doyeon Weon , Kyungchun Lee

Deep learning models have achieved remarkable success in different areas of machine learning over the past decade; however, the size and complexity of these models make them difficult to understand. In an effort to make them more…

计算机视觉与模式识别 · 计算机科学 2022-06-20 Vikram V. Ramaswamy , Sunnie S. Y. Kim , Nicole Meister , Ruth Fong , Olga Russakovsky

Prior knowledge about the imaging physics provides a mechanistic forward operator that plays an important role in image reconstruction, although myriad sources of possible errors in the operator could negatively impact the reconstruction…

图像与视频处理 · 电气工程与系统科学 2022-11-04 Maryam Toloubidokhti , Nilesh Kumar , Zhiyuan Li , Prashnna K. Gyawali , Brian Zenger , Wilson W. Good , Rob S. MacLeod , Linwei Wang

The linear spline growth model (LSGM), which approximates complex patterns using at least two linear segments, is a popular tool for examining nonlinear change patterns. Among such models, the linear-linear piecewise change pattern is the…

统计方法学 · 统计学 2022-05-10 Jin Liu , Robert A. Perera , Le Kang , Robert M. Kirkpatrick , Roy T. Sabo

This paper concerns the data-driven sensor deployment problem in large spatiotemporal fields. Traditionally, sensor deployment strategies have been heavily dependent on model-based planning approaches. However, model-based approaches do not…

信号处理 · 电气工程与系统科学 2022-01-04 Jiahong Chen

GPT is an auto-regressive Transformer-based pre-trained language model which has attracted a lot of attention in the natural language processing (NLP) domain due to its state-of-the-art performance in several downstream tasks. The success…

计算与语言 · 计算机科学 2021-10-18 Ali Edalati , Marzieh Tahaei , Ahmad Rashid , Vahid Partovi Nia , James J. Clark , Mehdi Rezagholizadeh

Mechanistic interpretability seeks to reverse engineer a trained neural network by identifying the minimal subset of internal components. We perform a mechanistic interpretability analysis of the Particle Transformer architecture, trained…

高能物理 - 唯象学 · 物理学 2026-05-12 Saurabh Rai , Sanmay Ganguly

The Gottesman-Kitaev-Preskill (GKP) error correcting code encodes a finite dimensional logical space in one or more bosonic modes, and has recently been demonstrated in trapped ions and superconducting microwave cavities. In this work we…

量子物理 · 物理学 2024-03-05 Mackenzie H. Shaw , Andrew C. Doherty , Arne L. Grimsmo

One of the unspoken challenges of tractography is choosing the right parameters for a given dataset or bundle. In order to tackle this challenge, we explore the multi-dimensional parameter space of tractography using streamline-specific…

图像与视频处理 · 电气工程与系统科学 2024-08-12 Ruben Vink , Anna Vilanova , Maxime Chamberland

Mechanistic interpretability has transformed the analysis of transformer circuits by decomposing model behavior into competing algorithms, identifying phase transitions during training, and deriving closed-form predictions for when and why…

机器学习 · 计算机科学 2026-03-19 Alma Lago

Gaussian processes (GPs) are a powerful tool for probabilistic inference over functions. They have been applied to both regression and non-linear dimensionality reduction, and offer desirable properties such as uncertainty estimates,…

机器学习 · 统计学 2014-10-01 Yarin Gal , Mark van der Wilk , Carl E. Rasmussen

Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize two representational hypotheses: (1) a representation in the…

Recent advances in fine-grained representation learning leverage local-to-global (emergent) relationships for achieving state-of-the-art results. The relational representations relied upon by such methods, however, are abstract. We aim to…

计算机视觉与模式识别 · 计算机科学 2023-10-25 Abhra Chaudhuri , Massimiliano Mancini , Zeynep Akata , Anjan Dutta

Mechanistic interpretability aims to explain neural model behaviour by reverse-engineering learned computational structure into human-understandable components. Without a formal framework, however, mechanistic explanations cannot be…

机器学习 · 计算机科学 2026-05-12 Ward Gauderis , Thomas Dooms , Steven T. Holmer , Kola Ayonrinde , Geraint A. Wiggins

For an explanation of a deep learning model to be effective, it must provide both insight into a model and suggest a corresponding action in order to achieve some objective. Too often, the litany of proposed explainable deep learning…

机器学习 · 计算机科学 2020-10-09 Laura Rieger , Chandan Singh , W. James Murdoch , Bin Yu