English
Related papers

Related papers: SigMA: Path Signatures and Multi-head Attention fo…

200 papers

To perform few-shot learning, language models extract signals from a few input-label pairs, aggregate these into a learned prediction rule, and apply this rule to new inputs. How is this implemented in the forward pass of modern transformer…

Machine Learning · Computer Science 2025-10-10 Xinyan Hu , Kayo Yin , Michael I. Jordan , Jacob Steinhardt , Lijie Chen

Gliomas are aggressive brain tumors that require accurate imaging-based diagnosis, with segmentation playing a critical role in evaluating morphology and treatment decisions. Manual delineation of gliomas is time-consuming and prone to…

Image and Video Processing · Electrical Eng. & Systems 2025-12-02 Cecilia Diana-Albelda , Roberto Alcover-Couso , Álvaro García-Martín , Jesus Bescos , Marcos Escudero-Viñolo

Softmax attention is the principle backbone of foundation models for various artificial intelligence applications, yet its quadratic complexity in sequence length can limit its inference throughput in long-context settings. To address this…

Machine Learning · Computer Science 2024-12-10 Jerome Sieber , Carmen Amo Alonso , Alexandre Didier , Melanie N. Zeilinger , Antonio Orvieto

This study established a feature-enhanced adversarial semi-supervised semantic segmentation model to automatically annotate pulmonary embolism lesion areas in computed tomography pulmonary angiogram (CTPA) images. In current studies, all of…

Image and Video Processing · Electrical Eng. & Systems 2022-04-12 Ting-Wei Cheng , Jerry Chang , Ching-Chun Huang , Chin Kuo , Yun-Chien Cheng

In this paper we construct a framework for doing statistical inference for discretely observed stochastic differential equations (SDEs) where the driving noise has 'memory'. Classical SDE models for inference assume the driving noise to be…

Methodology · Statistics 2013-07-05 Martin Lysy , Natesh S. Pillai

Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportunities for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-14 Alimjan Mattursun , Liejun Wang , Yinfeng Yu

Distribution Regression (DR) on stochastic processes describes the learning task of regression on collections of time series. Path signatures, a technique prevalent in stochastic analysis, have been used to solve the DR problem. Recent…

Machine Learning · Computer Science 2024-10-15 Andrew Alden , Carmine Ventre , Blanka Horvath

This study introduces a systematic framework to compare the efficacy of Large Language Models (LLMs) for fine-tuning across various cheminformatics tasks. Employing a uniform training methodology, we assessed three well-known…

Machine Learning · Computer Science 2024-05-03 Lee Youngmin , Lang S. I. D. Andrew , Cai Duoduo , Wheat R. Stephen

This paper addresses the challenges of mining latent patterns and modeling contextual dependencies in complex sequence data. A sequence pattern mining algorithm is proposed by integrating Bidirectional Long Short-Term Memory (BiLSTM) with a…

Machine Learning · Computer Science 2025-04-22 Tao Yang , Yu Cheng , Yaokun Ren , Yujia Lou , Minggu Wei , Honghui Xin

This paper investigates a non-autonomous slow-fast system, which is generalized by stochastic differential equations (SDEs) with locally Lipschitz coefficients, subjected to standard Brownian motion (Bm) and fractional Brownian motion (fBm)…

Probability · Mathematics 2020-12-21 Ruifang Wang , Yong Xu , Hongge Yue

The paradigm of Transformers using the self-attention mechanism has manifested its advantage in learning graph-structured data. Yet, Graph Transformers are capable of modeling full range dependencies but are often deficient in extracting…

Machine Learning · Computer Science 2024-09-11 Minhong Zhu , Zhenhao Zhao , Weiran Cai

We study the relationship between mixed stochastic differential equations and the corresponding rough path equations driven by standard Brownian motion and fractional Brownian motion with Hurst parameter $H>1/2$. We establish a correction…

Probability · Mathematics 2015-04-28 Andreas Neuenkirch , Taras Shalaiko

Transformer-based models have shown strong performance in speech deepfake detection, largely due to the effectiveness of the multi-head self-attention (MHSA) mechanism. MHSA provides frame-level attention scores, which are particularly…

Sound · Computer Science 2026-02-05 Tuan Dat Phuong , Duc-Tuan Truong , Long-Vu Hoang , Trang Nguyen Thi Thu

Stochastic differential equations (SDEs) provide a flexible framework for modeling temporal dynamics in partially observed systems. A central task is to calibrate such models from data, which requires inferring latent trajectories and…

Machine Learning · Statistics 2026-05-08 Yu Wang , Arnab Ganguly

Deep neural networks have revolutionized machine learning, yet their training dynamics remain theoretically unclear-we develop a continuous-time, matrix-valued stochastic differential equation (SDE) framework that rigorously connects the…

Machine Learning · Computer Science 2026-02-10 Brian Richard Olsen , Sam Fatehmanesh , Frank Xiao , Adarsh Kumarappan , Anirudh Gajula

Structural monitoring for complex built environments often suffers from mismatch between design, laboratory testing, and actual built parameters. Additionally, real-world structural identification problems encounter many challenges. For…

Machine Learning · Computer Science 2022-08-29 Xuyang Li , Hamed Bolandi , Talal Salem , Nizar Lajnef , Vishnu Naresh Boddeti

Multi-Head Attention (MHA) is a key component of Transformer. In MHA, attention heads work independently, causing problems such as low-rank bottleneck of attention score matrices and head redundancy. We propose Dynamically Composable…

Machine Learning · Computer Science 2024-06-05 Da Xiao , Qingye Meng , Shengping Li , Xingyuan Yuan

The reconstruction and inference of stochastic dynamical systems from data is a fundamental task in inverse problems and statistical learning. While surrogate modeling advances computational methods to approximate these dynamics, standard…

Optimization and Control · Mathematics 2026-04-14 Nicole Tianjiao Yang

Deep neural networks are typically trained in a single shot for a specific task and data distribution, but in real world settings both the task and the domain of application can change. The problem becomes even more challenging in dense…

Computer Vision and Pattern Recognition · Computer Science 2022-06-29 Donald Shenaj , Francesco Barbato , Umberto Michieli , Pietro Zanuttigh

We provide an introduction to the signature method, focusing on its theoretical properties and machine learning applications. Our presentation is divided into two parts. In the first part, we present the definition and fundamental…

Machine Learning · Statistics 2025-12-29 Ilya Chevyrev , Andrey Kormilitzin