中文
相关论文

相关论文: Sink vs. diagonal patterns as mechanisms for atten…

200 篇论文

This work aims to predict channels in wireless communication systems based on noisy observations, utilizing sequence-to-sequence models with attention (Seq2Seq-attn) and transformer models. Both models are adapted from natural language…

机器学习 · 统计学 2025-09-05 Valentina Rizzello , Benedikt Böck , Michael Joham , Wolfgang Utschick

Attention, specifically scaled dot-product attention, has proven effective for natural language, but it does not have a mechanism for handling hierarchical patterns of arbitrary nesting depth, which limits its ability to recognize certain…

计算与语言 · 计算机科学 2024-01-25 Brian DuSell , David Chiang

Using Monte Carlo techniques, we study a simple model which exhibits a competition between superconductivity and other types of order in two dimensions. The model is a site-diluted XY model, in which the XY spins are mobile, and also…

超导电性 · 物理学 2009-11-11 Daniel Valdez-Balderas , David Stroud

We formulate an attention mechanism for continuous and ordered sequences that explicitly functions as an alignment model, which serves as the core of many sequence-to-sequence tasks. Standard scaled dot-product attention relies on…

机器学习 · 计算机科学 2025-09-19 Hyungjoon Soh , Junghyo Jo

Word alignment, which aims to align translationally equivalent words between source and target sentences, plays an important role in many natural language processing tasks. Current unsupervised neural alignment methods focus on inducing…

计算与语言 · 计算机科学 2021-05-18 Chi Chen , Maosong Sun , Yang Liu

Understanding narratives requires identifying which events are most salient for a story's progression. We present a contrastive learning framework for modeling narrative salience that learns story embeddings from narrative twins: stories…

计算与语言 · 计算机科学 2026-01-13 Igor Sterner , Alex Lascarides , Frank Keller

Transformer models typically calculate attention matrices using dot products, which have limitations when capturing nonlinear relationships between embedding vectors. We propose Neural Attention, a technique that replaces dot products with…

机器学习 · 计算机科学 2025-11-10 Andrew DiGiugno , Ausif Mahmood

Building on recent advances in representation learning for wireless channels, this work investigates the cost-benefit trade-offs of high-dimensional channel embeddings in practical systems. We benchmark multiple wireless representations:…

信号处理 · 电气工程与系统科学 2026-05-05 Murilo Batista , Shirin Salehi , Saeed Mashdour , Paul Zheng , Rodrigo C. de Lamare , Anke Schmeink

Manifold models provide low-dimensional representations that are useful for processing and analyzing data in a transformation-invariant way. In this paper, we study the problem of learning smooth pattern transformation manifolds from image…

计算机视觉与模式识别 · 计算机科学 2013-05-20 Elif Vural , Pascal Frossard

Contrastive learning, along with its variations, has been a highly effective self-supervised learning method across diverse domains. Contrastive learning measures the distance between representations using cosine similarity and uses…

机器学习 · 计算机科学 2023-10-11 Daniel Rho , TaeSoo Kim , Sooill Park , Jaehyun Park , JaeHan Park

We consider the problem of predicting how the likelihood of an outcome of interest for a patient changes over time as we observe more of the patient data. To solve this problem, we propose a supervised contrastive learning framework that…

机器学习 · 计算机科学 2024-04-16 Shahriar Noroozizadeh , Jeremy C. Weiss , George H. Chen

Since their inception, CNNs have utilized some type of striding operator to reduce the overlap of receptive fields and spatial dimensions. Although having clear heuristic motivations (i.e. lowering the number of parameters to learn) the…

机器学习 · 计算机科学 2017-12-08 Chen Kong , Simon Lucey

The organization of latent token representations plays a crucial role in determining the stability, generalization, and contextual consistency of language models, yet conventional approaches to embedding refinement often rely on parameter…

计算与语言 · 计算机科学 2025-03-26 Meiquan Dong , Haoran Liu , Yan Huang , Zixuan Feng , Jianhong Tang , Ruoxi Wang

In this paper we consider a tank containing fluid and we want to estimate the horizontal currents when the fluid surface height is measured. The fluid motion is described by shallow water equations in two horizontal dimensions. We build a…

最优化与控制 · 数学 2013-09-20 Didier Auroux , S. Bonnabel

Standard attention-based transformers are known to exhibit instability under learning rate overspecification during training, particularly at high learning rates. While various methods have been proposed to improve resilience to such…

机器学习 · 计算机科学 2026-02-02 Shyam Venkatasubramanian , Sean Moushegian , Michael Lin , Mir Park , Ankit Singhal , Connor Lee

In recent years, hypergraph learning has attracted great attention due to its capacity in representing complex and high-order relationships. However, current neural network approaches designed for hypergraphs are mostly shallow, thus…

机器学习 · 计算机科学 2022-11-03 Guanzi Chen , Jiying Zhang , Xi Xiao , Yang Li

Dimensionality reduction is the essence of many data processing problems, including filtering, data compression, reduced-order modeling and pattern analysis. While traditionally tackled using linear tools in the fluid dynamics community,…

流体动力学 · 物理学 2023-02-01 Miguel A. Mendez

While static word embedding models are known to represent linguistic analogies as parallel lines in high-dimensional space, the underlying mechanism as to why they result in such geometric structures remains obscure. We find that an…

计算与语言 · 计算机科学 2023-06-16 Narutatsu Ri , Fei-Tzin Lee , Nakul Verma

Recent studies indicate that hierarchical Vision Transformer with a macro architecture of interleaved non-overlapped window-based self-attention \& shifted-window operation is able to achieve state-of-the-art performance in various visual…

计算机视觉与模式识别 · 计算机科学 2021-09-13 Yuxin Fang , Xinggang Wang , Rui Wu , Wenyu Liu

We characterize a prevalent weakness of deep neural networks (DNNs)---overthinking---which occurs when a DNN can reach correct predictions before its final layer. Overthinking is computationally wasteful, and it can also be destructive…

机器学习 · 计算机科学 2019-05-10 Yigitcan Kaya , Sanghyun Hong , Tudor Dumitras
‹ 上一页 1 8 9 10 下一页 ›