English
Related papers

Related papers: OWL: Geometry-Aware Spatial Reasoning for Audio La…

200 papers

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining minimal manual supervision with large-scale Visual Question-Answering (VQA) pairs…

Robotics · Computer Science 2026-05-05 Yangzhe Kong , Daeun Song , Jing Liang , Dinesh Manocha , Ziyu Yao , Xuesu Xiao

In the era of "Software Engineering 2.0" (SE 2.0), where intelligent agents collaborate with human engineers, Generative AI is advancing beyond code generation into Software Architecture (SA). While Large Language Models (LLMs) demonstrate…

Software Engineering · Computer Science 2026-03-10 Ha Vo , Nhut Tran , Khang Vo , Phat T. Tran-Truong , Son Ha

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Junfei Wu , Jian Guan , Kaituo Feng , Qiang Liu , Shu Wu , Liang Wang , Wei Wu , Tieniu Tan

Acoustic word embeddings (AWEs) are vector representations such that different acoustic exemplars of the same word are projected nearby in the embedding space. In addition to their use in speech technology applications such as spoken term…

Computation and Language · Computer Science 2023-01-10 Badr M. Abdullah , Dietrich Klakow

Sound can convey significant information for spatial reasoning in our daily lives. To endow deep networks with such ability, we address the challenge of dense indoor prediction with sound in both 2D and 3D via cross-modal knowledge…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Heeseung Yun , Joonil Na , Gunhee Kim

Recent advancements in large audio-language models (LALMs) have shown impressive capabilities in understanding and reasoning about audio and speech information. However, these models still face challenges, including hallucinating…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Chun-Yi Kuan , Hung-yi Lee

Spatial reasoning plays a vital role in both human cognition and machine intelligence, prompting new research into language models' (LMs) capabilities in this regard. However, existing benchmarks reveal shortcomings in evaluating…

Computation and Language · Computer Science 2024-05-27 Fangjun Li , David C. Hogg , Anthony G. Cohn

Spatial reasoning is a crucial component of both biological and artificial intelligence. In this work, we present a comprehensive study of the capability of current state-of-the-art large language models (LLMs) on spatial reasoning. To…

Computation and Language · Computer Science 2024-06-10 Md Imbesat Hassan Rizvi , Xiaodan Zhu , Iryna Gurevych

Large Audio Language Models (LALMs) have been widely applied in real-time scenarios, such as in-car assistants and online meeting comprehension. In practice, audio inputs are often corrupted by device and environmental noise, leading to…

Sound · Computer Science 2026-01-13 Yuanhe Zhang , Jiayu Tian , Yibo Zhang , Shilinlu Yan , Liang Lin , Zhenhong Zhou , Li Sun , Sen Su

Neural scaling laws offer valuable insights for designing robust sequence processing architectures. While these laws have been extensively characterized in other modalities, their behavior in speech remains comparatively underexplored. In…

Computation and Language · Computer Science 2025-02-17 William Chen , Jinchuan Tian , Yifan Peng , Brian Yan , Chao-Han Huck Yang , Shinji Watanabe

Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural language instructions. However, when applied to robotics, VLMs remain subject to a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Xiaowen Sun , Matthias Kerzel , Mengdi Li , Xufeng Zhao , Paul Striker , Stefan Wermter

A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial…

Embodied AI aims to develop robots that can \textit{understand} and execute human language instructions, as well as communicate in natural languages. On this front, we study the task of generating highly detailed navigational instructions…

Computation and Language · Computer Science 2024-09-10 Muraleekrishna Gopinathan , Martin Masek , Jumana Abu-Khalaf , David Suter

We present Spatial LibriSpeech, a spatial audio dataset with over 650 hours of 19-channel audio, first-order ambisonics, and optional distractor noise. Spatial LibriSpeech is designed for machine learning model training, and it includes…

Answering real-world geospatial questions--such as finding restaurants along a travel route or amenities near a landmark--requires reasoning over both geographic relationships and semantic user intent. However, existing large language…

Information Retrieval · Computer Science 2025-06-12 Dazhou Yu , Riyang Bao , Ruiyu Ning , Jinghong Peng , Gengchen Mai , Liang Zhao

Audio large language models (ALLMs) have recently advanced spoken interaction by integrating speech processing with large language models. However, existing evaluations of fairness, safety, and security (FSS) remain fragmented, largely…

Sound · Computer Science 2026-03-17 Ranya Aloufi , Srishti Gupta , Soumya Shaw , Battista Biggio , Lea Schönherr

While Large Audio Language Models (LALMs) achieve strong performance on short audio, they degrade on long-form inputs. This degradation is more severe in temporal awareness tasks, where temporal alignment becomes increasingly inaccurate as…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Mingchen Shao , Hang Su , Wenjie Tian , Bingshen Mu , Zhennan Lin , Lichun Fan , Zhenbo Luo , Jian Luan , Lei Xie

With the rapid development of spatial audio technologies today, applications in AR, VR, and other scenarios have garnered extensive attention. Unlike traditional mono sound, spatial audio offers a more realistic and immersive auditory…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Zhiyuan Zhu , Yu Zhang , Wenxiang Guo , Changhao Pan , Zhou Zhao

Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remains unreliable, and existing approaches largely rely on data…

Sound · Computer Science 2026-02-12 Liyang Chen , Hongkai Chen , Yujun Cai , Sifan Li , Qingwen Ye , Yiwei Wang

Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental exploration. As the intrinsic properties of the environment determine the abilities that…

Computation and Language · Computer Science 2026-05-04 Peng Yu , Zeyuan Zhao , Shao Zhang , Luoyi Fu , Xinbing Wang , Ying Wen