English
Related papers

Related papers: AIBA: Attention-based Instrument Band Alignment fo…

200 papers

In millimeter-wave (mmWave) communications, directional transmission based on beamforming is important to compensate for high pathloss. To maintain the desired direction transmission gain, beam scanning that involves the transmitter sending…

Systems and Control · Electrical Eng. & Systems 2024-01-02 Huang-Chou Lin , Kuang-Hao , Liu

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific…

Recent Transformer-based diffusion models have shown remarkable performance, largely attributed to the ability of the self-attention mechanism to accurately capture both global and local contexts by computing all-pair interactions among…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Yunxiang Fu , Chaoqi Chen , Yizhou Yu

In recent years, there has been a growing emphasis on the intersection of audio, vision, and text modalities, driving forward the advancements in multimodal research. However, strong bias that exists in any modality can lead to the model…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Xiulong Liu , Zhikang Dong , Peng Zhang

Recent advancements in diffusion models have notably improved the perceptual quality of generated images in text-to-image synthesis tasks. However, diffusion models often struggle to produce images that accurately reflect the intended…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Yang Zhang , Teoh Tze Tzun , Lim Wei Hern , Tiviatis Sim , Kenji Kawaguchi

Recent text-to-image (T2I) generation models have advanced significantly, enabling the creation of high-fidelity images from textual prompts. However, existing evaluation benchmarks primarily focus on the explicit alignment between…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Wenchao Zhang , Jiahe Tian , Runze He , Jizhong Han , Jiao Dai , Miaomiao Feng , Wei Mi , Xiaodan Zhang

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Eitan Shaar , Ariel Shaulov , Gal Chechik , Lior Wolf

The Laser Interferometer Space Antenna (LISA) mission features a three-spacecraft long-arm constellation intended to detect gravitational wave sources in the low-frequency band up to 1 Hz via laser interferometry. The paper presents an…

Instrumentation and Detectors · Physics 2022-08-24 Niklas Houba , Simon Delchambre , Tobias Ziegler , Gerald Hechenblaikner , Walter Fichter

Attention regulates information transfer between tokens. For this, query and key vectors are compared, typically in terms of a scalar product, $\mathbf{Q}^T\mathbf{K}$, together with a subsequent softmax normalization. In geometric terms,…

Machine Learning · Computer Science 2025-08-11 Claudius Gros

This paper introduces a novel combination of two tasks, previously treated separately: acoustic-to-articulatory speech inversion (AAI) and phoneme-to-articulatory (PTA) motion estimation. We refer to this joint task as acoustic…

Cross-attention is the primary interface through which text conditions latent diffusion models, yet its step-wise multi-resolution dynamics remain under-characterized, limiting principled training-free control. We cast diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Seunghun Oh , Unsang Park

We propose Adaptive Integrated Layered Attention (AILA), a neural network architecture that combines dense skip connections with different mechanisms for adaptive feature reuse across network layers. We evaluate AILA on three challenging…

Machine Learning · Computer Science 2025-05-14 William Claster , Suhas KM , Dhairya Gundechia

Large Audio Language Models (LALMs) demonstrate impressive general audio understanding, but once deployed, they are static and fail to improve with new real-world audio data. As traditional supervised fine-tuning is costly, we introduce a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Haoyu Zhang , Jiaxian Guo , Yusuke Iwasawa , Yutaka Matsuo

Although attention mechanisms have achieved considerable progress in Transformer-based architectures across various Artificial Intelligence (AI) domains, their inner workings remain to be explored. Existing explainable methods have…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Hongbo Zhu , Theodor Wulff , Rahul Singh Maharjan , Jinpei Han , Angelo Cangelosi

The two-pass information bottleneck (TPIB) based speaker diarization system operates independently on different conversational recordings. TPIB system does not consider previously learned speaker discriminative information while diarizing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-14 Nauman Dawalatabad , Srikanth Madikeri , C Chandra Sekhar , Hema A Murthy

Unsupervised domain adaptation (UDA) for person re-identification is challenging because of the huge gap between the source and target domain. A typical self-training method is to use pseudo-labels generated by clustering algorithms to…

Computer Vision and Pattern Recognition · Computer Science 2021-12-28 Wenhao Wang , Fang Zhao , Shengcai Liao , Ling Shao

Audio representation learning typically evaluates design choices such as input frontend, sequence backbone, and sequence length in isolation. We show that these axes are coupled, and conclusions from one setting often do not transfer to…

Sound · Computer Science 2026-03-24 Khushiyant , Param Thakkar

Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Kaixiang Yang , Boyang Shen , Xin Li , Yuchen Dai , Yuxuan Luo , Yueran Ma , Wei Fang , Qiang Li , Zhiwei Wang

Audio question answering (AQA) is the task of producing natural language answers when a system is provided with audio and natural language questions. In this paper, we propose neural network architectures based on self-attention and…

Computation and Language · Computer Science 2023-06-01 Parthasaarathy Sudarsanam , Tuomas Virtanen

The feature attribution method reveals the contribution of input variables to the decision-making process to provide an attribution map for explanation. Existing methods grounded on the information bottleneck principle compute information…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jung-Ho Hong , Ho-Joong Kim , Kyu-Sung Jeon , Seong-Whan Lee