English
Related papers

Related papers: Can Large Language Models Understand Spatial Audio…

200 papers

The revolutionary capabilities of large language models (LLMs) have paved the way for multimodal large language models (MLLMs) and fostered diverse applications across various specialized domains. In the remote sensing (RS) field, however,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Dilxat Muhtar , Zhenshi Li , Feng Gu , Xueliang Zhang , Pengfeng Xiao

Although state-of-the-art Speech Foundational Models can produce high-quality text pseudo-labels, applying Semi-Supervised Learning (SSL) for in-the-wild real-world data remains challenging due to its richer and more complex acoustics…

Computation and Language · Computer Science 2026-03-16 Wen Ding , Fan Qian

Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their application to remote sensing (RS) remains underexplored due to significant domain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Sungjune Park , Yeongyun Kim , Se Yeon Kim , Yong Man Ro

Speaker Diarization (SD) is a crucial component of modern end-to-end ASR pipelines. Traditional SD systems, which are typically audio-based and operate independently of ASR, often introduce speaker errors, particularly during speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Anurag Kumar , Rohit Paturi , Amber Afshan , Sundararajan Srinivasan

Understanding the meaning of words in context is a fundamental capability for Large Language Models (LLMs). Despite extensive evaluation efforts, the extent to which LLMs show evidence that they truly grasp word senses remains…

Computation and Language · Computer Science 2025-09-18 Domenico Meconi , Simone Stirpe , Federico Martelli , Leonardo Lavalle , Roberto Navigli

Audio Large Language Models (AudioLLMs) have received widespread attention and have significantly improved performance on audio tasks such as conversation, audio understanding, and automatic speech recognition (ASR). Despite these…

Computational Engineering, Finance, and Science · Computer Science 2025-12-19 Yupeng Cao , Haohang Li , Yangyang Yu , Shashidhar Reddy Javaji , Yueru He , Jimin Huang , Qianqian Xie , Fabrizio Dimino , Xiao-yang Liu , K. P. Subbalakshmi , Meikang Qiu , Sophia Ananiadou , Jian-Yun Nie

Multimodal Large Language Models (MLLMs) that directly process RGB inputs for tasks like 3D localization and navigation have shown remarkable potential. However, we argue that these RGB-only approaches are fundamentally flawed in their…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Gongjie Zhang , Wenhao Li , Quanhao Qian , Jiuniu Wang , Deli Zhao , Shijian Lu , Ran Xu

With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for certain tasks without supervised data. We extend this…

Computation and Language · Computer Science 2026-01-29 Zhaolin Li , Jan Niehues

An ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs…

Spatial understanding has been a challenging task for existing Multi-modal Large Language Models~(MLLMs). Previous methods leverage large-scale MLLM finetuning to enhance MLLM's spatial understanding ability. In this paper, we present a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Hsiang-Wei Huang , Jen-Hao Cheng , Kuang-Ming Chen , Cheng-Yen Yang , Bahaa Alattar , Yi-Ru Lin , Pyongkun Kim , Sangwon Kim , Kwangju Kim , Chung-I Huang , Jenq-Neng Hwang

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR,…

Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) \etc., which can be challenging to cover using one small model designed for a single task. In…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-02 Siyin Wang , Wenyi Yu , Yudong Yang , Changli Tang , Yixuan Li , Jimin Zhuang , Xianzhao Chen , Xiaohai Tian , Jun Zhang , Guangzhi Sun , Lu Lu , Yuxuan Wang , Chao Zhang

Spatial perception is central to auditory intelligence, enabling accurate understanding of real-world acoustic scenes and advancing human-level perception of the world around us. While recent large audio-language models (LALMs) show strong…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-17 S Sakshi , Vaibhavi Lokegaonkar , Neil Zhang , Ramani Duraiswami , Sreyan Ghosh , Dinesh Manocha , Lie Lu

We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehension, they still…

Computation and Language · Computer Science 2025-10-23 Yejin Kwon , Taewoo Kang , Hyunsoo Yoon , Changouk Kim

Recent works have shown that Deep Recurrent Neural Networks using the LSTM architecture can achieve strong single-channel speech enhancement by estimating time-frequency masks. However, these models do not naturally generalize to…

Sound · Computer Science 2020-12-04 Felix Grezes , Zhaoheng Ni , Viet Anh Trinh , Michael Mandel

Large language models (LLMs) have recently achieved impressive results in speech recognition across multiple modalities, including Auditory Speech Recognition (ASR), Visual Speech Recognition (VSR), and Audio-Visual Speech Recognition…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Umberto Cappellazzo , Xubo Liu , Pingchuan Ma , Stavros Petridis , Maja Pantic

We propose to utilize an instruction-tuned large language model (LLM) for guiding the text generation process in automatic speech recognition (ASR). Modern large language models (LLMs) are adept at performing various text generation tasks…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-08 Yosuke Higuchi , Tetsuji Ogawa , Tetsunori Kobayashi

Recently, Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated aptitude as potential substitutes for human participants in experiments testing psycholinguistic phenomena. However, an understudied question is to…

Computation and Language · Computer Science 2024-10-21 Tyler Loakman , Yucheng Li , Chenghua Lin

The capabilities of large language models (LLMs) have sparked debate over whether such systems just learn an enormous collection of superficial statistics or a set of more coherent and grounded representations that reflect the real world.…

Machine Learning · Computer Science 2024-03-05 Wes Gurnee , Max Tegmark

Large language models (LLMs) are rapidly transforming materials science. This review examines recent LLM applications across the materials discovery pipeline, focusing on three key areas: mining scientific literature , predictive modelling,…

Computation and Language · Computer Science 2025-11-17 Fengxu Yang , Weitong Chen , Jack D. Evans