English
Related papers

Related papers: AMAuT: A Flexible and Efficient Multiview Audio Tr…

200 papers

Traditional Automated Speaking Assessment (ASA) systems exhibit inherent modality limitations: text-based approaches lack acoustic information while audio-based methods miss semantic context. Multimodal Large Language Models (MLLM) offer…

Computation and Language · Computer Science 2025-08-19 Yu-Hsuan Fang , Tien-Hong Lo , Yao-Ting Sung , Berlin Chen

AudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and completeness remain critical bottlenecks that limit…

Sound · Computer Science 2025-08-25 Yulin Sun , Qisheng Xu , Yi Su , Qian Zhu , Yong Dou , Xinwang Liu , Kele Xu

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a challenge, in part…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Edson Araujo , Saurabhchand Bhati , M. Jehanzeb Mirza , Brian Kingsbury , Samuel Thomas , Rogerio Feris , James R. Glass , Hilde Kuehne

Large Audio Language Models (LALMs) demonstrate impressive general audio understanding, but once deployed, they are static and fail to improve with new real-world audio data. As traditional supervised fine-tuning is costly, we introduce a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Haoyu Zhang , Jiaxian Guo , Yusuke Iwasawa , Yutaka Matsuo

Speech quality and intelligibility are significantly degraded in noisy environments. This paper presents a novel transformer-based learning framework to address the single-channel noise suppression problem for real-time applications.…

Sound · Computer Science 2025-11-18 Behnaz Bahmei , Siamak Arzanpour , Elina Birmingham

Consonant and vowel reduction are often encountered in speech, which might cause performance degradation in automatic speech recognition (ASR). Our recently proposed learning strategy based on masking, Phone Masking Training (PMT),…

Sound · Computer Science 2022-07-05 Guodong Ma , Pengfei Hu , Nurmemet Yolwas , Shen Huang , Hao Huang

Transformers (Vaswani et al., 2017) have brought a remarkable improvement in the performance of neural machine translation (NMT) systems but they could be surprisingly vulnerable to noise. In this work, we try to investigate how noise…

Computation and Language · Computer Science 2021-09-13 Peyman Passban , Puneeth S. M. Saladi , Qun Liu

Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as shifting the order of choices, result in substantially…

Computation and Language · Computer Science 2025-10-07 Fernando López , Santosh Kesiraju , Jordi Luque

A key challenge in machine learning is to generalize from training data to an application domain of interest. This work generalizes the recently-proposed mixture invariant training (MixIT) algorithm to perform unsupervised learning in the…

Sound · Computer Science 2024-03-25 Cong Han , Kevin Wilson , Scott Wisdom , John R. Hershey

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval,…

Sound · Computer Science 2026-01-30 Hui Wang , Jinghua Zhao , Junyang Cheng , Cheng Liu , Yuhang Jia , Haoqin Sun , Jiaming Zhou , Yong Qin

Audio classifiers frequently face domain shift, when models trained on one dataset lose accuracy on data recorded in acoustically different conditions. Previous Test-Time Adaptation (TTA) research in speech and sound analysis often…

Sound · Computer Science 2025-11-25 Weichuang Shao , Iman Yi Liao , Tomas Henrique Bode Maul , Tissa Chandesa

Recent deep multi-view stereo (MVS) methods have widely incorporated transformers into cascade network for high-resolution depth estimation, achieving impressive results. However, existing transformer-based methods are constrained by their…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Sicheng Wang , Hao Jiang , Lei Xiang

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a…

Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions frequently lead to improvements in performance and may even be critical in certain cases, they come…

Sound · Computer Science 2025-01-31 Joanna Hong , Sanjeel Parekh , Honglie Chen , Jacob Donley , Ke Tan , Buye Xu , Anurag Kumar

Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks that require structured…

Sound · Computer Science 2025-10-14 Kuan-Yi Lee , Tsung-En Lin , Hung-Yi Lee

This paper presents a way of doing large scale audio understanding without traditional state of the art neural architectures. Ever since the introduction of deep learning for understanding audio signals in the past decade, convolutional…

Sound · Computer Science 2022-02-01 Prateek Verma

This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large Language Models (MLLMs) can respond to audio-visual content,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-08 Qilang Ye , Zitong Yu , Rui Shao , Xinyu Xie , Philip Torr , Xiaochun Cao

In line with the human capacity to perceive the world by simultaneously processing and integrating high-dimensional inputs from multiple modalities like vision and audio, we propose a novel model, MAiVAR-T (Multimodal Audio-Image to Video…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Muhammad Bilal Shaikh , Douglas Chai , Syed Mohammed Shamsul Islam , Naveed Akhtar

In this paper, we aim to unveil the impact of data augmentation in audio-language multi-modal learning, which has not been explored despite its importance. We explore various augmentation methods at not only train-time but also test-time…

Sound · Computer Science 2023-05-24 Eungbeom Kim , Jinhee Kim , Yoori Oh , Kyungsu Kim , Minju Park , Jaeheon Sim , Jinwoo Lee , Kyogu Lee

Considering the bimodal nature of human speech perception, lips, and teeth movement has a pivotal role in automatic speech recognition. Benefiting from the correlated and noise-invariant visual information, audio-visual recognition systems…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-23 Xiaoming Ren , Chao Li , Shenjian Wang , Biao Li