English
Related papers

Related papers: AIBA: Attention-based Instrument Band Alignment fo…

200 papers

Recent text-to-image (T2I) diffusion models have achieved remarkable advancement, yet faithfully following complex textual descriptions remains challenging due to insufficient interactions between textual and visual features. Prior…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Binglei Li , Mengping Yang , Zhiyu Tan , Junping Zhang , Hao Li

Vision transformer has demonstrated great potential in abundant vision tasks. However, it also inevitably suffers from poor generalization capability when the distribution shift occurs in testing (i.e., out-of-distribution data). To…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Xin Li , Cuiling Lan , Guoqiang Wei , Zhibo Chen

Audio tagging is an important task of mapping audio samples to their corresponding categories. Recently endeavours that exploit transformer models in this field have achieved great success. However, the quadratic self-attention cost limits…

Sound · Computer Science 2024-05-24 Jiaju Lin , Haoxuan Hu

Driver gaze estimation serves as a fundamental metric for evaluating driver attentiveness in modern monitoring systems. Beyond being vulnerable to sudden lighting changes and sensor noise, spatial-domain models struggle to disentangle…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Jun Ma , Zhenye Yang , Ruichen Zhou , Pei Zhang , Huan Li , Jinpeng Chen

The direction-of-arrival (DOA) of sound sources is an essential acoustic parameter used, e.g., for multi-channel speech enhancement or source tracking. Complex acoustic scenarios consisting of sources-of-interest, interfering sources,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-17 Wolfgang Mack , Julian Wechsler , Emanuël A. P. Habets

Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottleneck. In this work, we propose In-context Sparse Attention (ISA), the first…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Shitong Shao , Zikai Zhou , Haopeng Li , Yingwei Song , Wenliang Zhong , Lichen Bai , Zeke Xie

While modern text-to-image models excel at prompt-based generation, they often lack the fine-grained control necessary for specific user requirements like spatial layouts or subject appearances. Multi-condition control addresses this, yet…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Chao Zhou , Tianyi Wei , Yiling Chen , Wenbo Zhou , Nenghai Yu

In complex auditory environments, the human auditory system possesses the remarkable ability to focus on a specific speaker while disregarding others. In this study, a new model named SWIM, a short-window convolution neural network (CNN)…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Ziyang Zhang , Andrew Thwaites , Alexandra Woolgar , Brian Moore , Chao Zhang

Preference alignment has emerged as an effective strategy to enhance the performance of Multimodal Large Language Models (MLLMs) following supervised fine-tuning. While existing preference alignment methods predominantly target…

Computer Vision and Pattern Recognition · Computer Science 2025-09-08 Zitian Wang , Yue Liao , Kang Rong , Fengyun Rao , Yibo Yang , Si Liu

Domain adaptation deals with training models using large scale labeled data from a specific source domain and then adapting the knowledge to certain target domains that have few or no labels. Many prior works learn domain agnostic feature…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Astuti Sharma , Tarun Kalluri , Manmohan Chandraker

The International Phonetic Alphabet (IPA) is indispensable in language learning and understanding, aiding users in accurate pronunciation and comprehension. Additionally, it plays a pivotal role in speech therapy, linguistic research,…

Computation and Language · Computer Science 2023-11-08 Jakir Hasan , Shrestha Datta , Ameya Debnath

Explaining the behavior of end-to-end audio language models via Shapley value attribution is intractable under native tokenization: a typical utterance yields over $150$ encoder frames, inflating the coalition space by roughly $10^{42}$…

Sound · Computer Science 2026-03-04 Paweł Pozorski , Jakub Muszyński , Maria Ganzha

Many tool-based Retrieval Augmented Generation (RAG) systems lack precise mechanisms for tracing final responses back to specific tool components -- a critical gap as systems scale to complex multi-agent architectures. We present…

Information Retrieval · Computer Science 2026-02-06 James Gao , Josh Zhou , Qi Sun , Ryan Huang , Steven Yoo

Ion Beam Analysis (IBA) comprises a set of analytical techniques suited for material analysis, many of which are rather closely related. Self-consistent analysis of several IBA techniques takes advantage of this close relationship to…

Domain Adaptation of Black-box Predictors (DABP) aims to learn a model on an unlabeled target domain supervised by a black-box predictor trained on a source domain. It does not require access to both the source-domain data and the predictor…

Machine Learning · Computer Science 2022-05-31 Jianfei Yang , Xiangyu Peng , Kai Wang , Zheng Zhu , Jiashi Feng , Lihua Xie , Yang You

Recent advancements in deep generative models present new opportunities for music production but also pose challenges, such as high computational demands and limited audio quality. Moreover, current systems frequently rely solely on text…

Sound · Computer Science 2024-10-31 Javier Nistal , Marco Pasini , Cyran Aouameur , Maarten Grachten , Stefan Lattner

Vision-Language Models (VLMs) have become essential backbones of modern multimodal intelligence, yet their outputs remain prone to hallucination-plausible text misaligned with visual inputs. Existing alignment approaches often rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Kejia Chen , Jiawen Zhang , Jiacong Hu , Kewei Gao , Jian Lou , Zunlei Feng , Mingli Song

Recent studies have revealed that text-to-image diffusion models are vulnerable to backdoor attacks, where attackers implant stealthy textual triggers to manipulate model outputs. Previous backdoor detection methods primarily focus on the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zhongqi Wang , Jie Zhang , Shiguang Shan , Xilin Chen

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

Connecting large libraries of digitized audio recordings to their corresponding sheet music images has long been a motivation for researchers to develop new cross-modal retrieval systems. In recent years, retrieval systems based on…

Information Retrieval · Computer Science 2019-06-27 Stefan Balke , Matthias Dorfer , Luis Carvalho , Andreas Arzt , Gerhard Widmer