English
Related papers

Related papers: SPUR: A Plug-and-Play Framework for Integrating Sp…

200 papers

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

Multimedia · Computer Science 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim

Spatial audio understanding aims to enable machines to interpret complex auditory scenes, particularly when sound sources move over time. In this work, we study Spatial Audio Question Answering (Spatial AQA) with a focus on movement…

Sound · Computer Science 2026-02-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

Current audio-visual (AV) benchmarks focus on final answer accuracy, overlooking the underlying reasoning process. This makes it difficult to distinguish genuine comprehension from correct answers derived through flawed reasoning or…

Computer Vision and Pattern Recognition · Computer Science 2025-08-22 Siminfar Samakoush Galougah , Rishie Raj , Sanjoy Chowdhury , Sayan Nag , Ramani Duraiswami

Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate LALMs' audio understanding performance, researchers…

Sound · Computer Science 2026-02-16 Han Yin , Jung-Woo Choi

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Haoyu Zhang , Meng Liu , Zaijing Li , Haokun Wen , Weili Guan , Yaowei Wang , Liqiang Nie

Acoustic scene perception involves describing the type of sounds, their timing, their direction and distance, as well as their loudness and reverberation. While audio language models excel in sound recognition, single-channel input…

Sound · Computer Science 2025-10-08 Xilin Jiang , Hannes Gamper , Sebastian Braun

Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-26 Chih-Kai Yang , Neo Ho , Yen-Ting Piao , Hung-yi Lee

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

Graphics · Computer Science 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Current vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessary for human-like understanding and real-world applications.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Wenyu Zhang , Wei En Ng , Lixin Ma , Yuwen Wang , Junqi Zhao , Allison Koenecke , Boyang Li , Lu Wang

Audio quality assessment is critical for assessing the perceptual realism of sounds. However, the time and expense of obtaining ''gold standard'' human judgments limit the availability of such data. For AR&VR, good perceived sound quality…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-27 Pranay Manocha , Anurag Kumar , Buye Xu , Anjali Menon , Israel D. Gebru , Vamsi K. Ithapu , Paul Calamia

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. However, most advanced…

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific…

Large Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio. However, most existing models primarily focus on analyzing input signals using text…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-20 Junyi Ao , Dekun Chen , Xiaohai Tian , Wenjie Feng , Jun Zhang , Lu Lu , Yuxuan Wang , Haizhou Li , Zhizheng Wu

The spatial sensitivity of an ultrasound transducer, which strongly influences its suitability for different applications, depends on the shape of the transducer surface. Accurate simulation of these spatial effects is important for…

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Yang Liu , Ming Ma , Xiaomin Yu , Pengxiang Ding , Han Zhao , Mingyang Sun , Siteng Huang , Donglin Wang

Large-scale pre-trained image-text models exhibit robust multimodal representations, yet applying the Contrastive Language-Image Pre-training (CLIP) model to audio-visual localization remains challenging. Replacing the classification token…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Khanh Binh Nguyen , Chae Jung Park

Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational…

The auditory system plays a substantial role in shaping the overall human perceptual experience. While prevailing large language models (LLMs) and visual language models (VLMs) have shown their promise in solving a wide variety of language…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-19 Jinhua Liang , Xubo Liu , Wenwu Wang , Mark D. Plumbley , Huy Phan , Emmanouil Benetos

As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization,…

Sound · Computer Science 2026-05-15 KiHyun Nam , Jungwoo Heo , Siu Bae , Ha-Jin Yu , Joon Son Chung

The ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications. Although significant progress has been made in this area since the development of AudioSet, most existing models…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-21 Yuan Gong , Hongyin Luo , Alexander H. Liu , Leonid Karlinsky , James Glass