English
Related papers

Related papers: OWL: Geometry-Aware Spatial Reasoning for Audio La…

200 papers

Large language models (LLMs) can generate syntactically valid optimization programs, yet often struggle to reliably choose an effective modeling strategy, leading to incorrect formulations and inefficient solver behavior. We propose SAGE, a…

Artificial Intelligence · Computer Science 2026-05-05 Ruiqing Zhao , Fengzhi Li , Yuan Zuo , Rui Liu , Yansong Liu , Yunfei Ma , Fanyu Meng , Junlan Feng

Despite the significant improvements achieved by large language models (LLMs) in English reasoning tasks, these models continue to struggle with multilingual reasoning. Recent studies leverage a full-parameter and two-stage training…

Computation and Language · Computer Science 2025-01-08 Yuchun Fan , Yongyu Mu , Yilin Wang , Lei Huang , Junhao Ruan , Bei Li , Tong Xiao , Shujian Huang , Xiaocheng Feng , Jingbo Zhu

Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that…

During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not…

Sound · Computer Science 2022-07-22 Wei Sun , Mei Wang , Lili Qiu

Spatial reasoning over three-dimensional scenes is a core capability for embodied intelligence, yet continuous model improvement remains bottlenecked by the cost of geometric annotation. The self-evolving paradigm offers a promising path,…

Large language models (LLMs) often generate reasoning traces that appear coherent but rest on unsupported assumptions, leading to hallucinated conclusions. Prior work mainly addresses factual hallucinations or relies on post-hoc…

Computation and Language · Computer Science 2025-10-21 Samir Abdaljalil , Erchin Serpedin , Khalid Qaraqe , Hasan Kurban

Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on…

Sound · Computer Science 2026-01-28 Congyi Fan , Jian Guan , Youtian Lin , Dongli Xu , Tong Ye , Qiaoxi Zhu , Pengming Feng , Wenwu Wang

While recent advances in Multimodal Large Language Models (MLLMs) have improved image-based localization, precise global geolocation remains a formidable challenge due to the inherent ambiguity of visual landscapes and the largely untapped…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Yiyang Su , Xiaoming Liu

Fine-grained spatiotemporal reasoning on surgical videos is critical, yet the capabilities of Multi-modal Large Language Models (MLLMs) in this domain remain largely unexplored. To bridge this gap, we introduce SurgCoT, a unified benchmark…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Gui Wang , YongSong Zhou , Kaijun Deng , Wooi Ping Cheah , Rong Qu , Jianfeng Ren , Linlin Shen

This paper studies audio-visual deep saliency prediction. It introduces a conceptually simple and effective Deep Audio-Visual Embedding for dynamic saliency prediction dubbed ``DAVE" in conjunction with our efforts towards building an…

Computer Vision and Pattern Recognition · Computer Science 2020-01-09 Hamed R. Tavakoli , Ali Borji , Esa Rahtu , Juho Kannala

Audio-Visual Navigation (AVN) requires an embodied agent to navigate toward a sound source by utilizing both vision and binaural audio. A core challenge arises in complex acoustic environments, where binaural cues become intermittently…

Sound · Computer Science 2026-04-06 Teng Liu , Yinfeng Yu

Multimodal large language models (MLLMs) have achieved strong performance on perception-oriented tasks, yet their ability to perform mathematical spatial reasoning, defined as the capacity to parse and manipulate two- and three-dimensional…

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced multimodal reasoning.…

Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize…

Computation and Language · Computer Science 2025-08-26 Chih-Kai Yang , Neo Ho , Yi-Jyun Lee , Hung-yi Lee

Contact-rich robotic manipulation requires representations that encode local geometry. Vision provides global context but lacks direct measurements of properties such as texture and hardness, whereas touch supplies these cues. Modern…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Gurmeher Khurana , Lan Wei , Dandan Zhang

Speech encoding models use auditory representations to predict how the human brain responds to spoken language stimuli. Most performant encoding models linearly map the hidden states of artificial neural networks to brain data, but this…

Computation and Language · Computer Science 2025-02-14 Nishitha Vattikonda , Aditya R. Vaidya , Richard J. Antonello , Alexander G. Huth

Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-04 Alan Pawlak , Hyunkook Lee , Aki Mäkivirta , Thomas Lund

Over the past few decades, computational methods have been developed to estimate perceptual audio quality. These methods, also referred to as objective quality measures, are usually developed and intended for a specific application domain.…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-25 Matteo Torcoli , Thorsten Kastner , Jürgen Herre

Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial audio models, conversely, are constrained to fixed microphone geometries, preventing deployment across…

Sound · Computer Science 2026-01-30 Artem Dementyev , Wazeer Zulfikar , Sinan Hersek , Pascal Getreuer , Anurag Kumar , Vivek Kumar

Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where beliefs must be continuously revised from egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chih-Ting Liao , Xi Xiao , Chunlei Meng , Zhangquan Chen , Yitong Qiao , Weilin Zhou , Tianyang Wang , Xu Zheng , Xin Cao