English
Related papers

Related papers: SPUR: A Plug-and-Play Framework for Integrating Sp…

200 papers

The sweet spot can be interpreted as the region where acoustic sources create a spatial auditory illusion. We study the problem of maximizing this sweet spot when reproducing a desired sound wave using an array of loudspeakers. To achieve…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-05 Pedro Izquierdo Lehmann , Rodrigo F. Cadiz , Carlos A. Sing Long

The foundational capabilities established by Large Language Models (LLMs) have paved the way for Multimodal Large Language Models (MLLMs), within which Large Audio Language Models (LALMs) are essential for realizing universal auditory…

A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved…

Spoken Language Understanding (SLU), which aims to extract user semantics to execute downstream tasks, is a crucial component of task-oriented dialog systems. Existing SLU datasets generally lack sufficient diversity and complexity, and…

Computation and Language · Computer Science 2025-12-02 Yuezhang Peng , Chonghao Cai , Ziang Liu , Shuai Fan , Sheng Jiang , Hua Xu , Yuxin Liu , Qiguang Chen , Kele Xu , Yao Li , Sheng Wang , Libo Qin , Xie Chen

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temporally consistent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Mohamad Alansari , Naufal Suryanto , Divya Velayudhan , Sajid Javed , Naoufel Werghi , Muzammal Naseer

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker,…

Sound · Computer Science 2025-12-24 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

In time-critical eXtended reality (XR) scenarios where users must rapidly reorient their attention to hazards, alerts, or instructions while engaged in a primary task, spatial audio can provide an immediate directional cue without occupying…

Human-Computer Interaction · Computer Science 2026-05-08 Yoonsang Kim , Swapnil Dey , Arie Kaufman

Long-form audio understanding poses significant challenges for large audio language models (LALMs) due to the extreme length of audio sequences and the need to reason over heterogeneous acoustic cues distributed over time, such as speech…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Masao Someki , Chien-yu Huang , Siddhant Arora , Samuele Cornell , Markus Müller , Nathan Susanj , Rupak V Swaminathan , Grant P Strimel , Jing Liu , Shinji Watanabe

Spatio-temporal forecasting plays a crucial role in various sectors such as transportation systems, logistics, and supply chain management. However, existing methods are limited by their ability to handle large, complex datasets. To…

Machine Learning · Computer Science 2024-08-27 Sakhinana Sagar Srinivas , Chidaksh Ravuru , Geethan Sannidhi , Venkataramana Runkana

LSTMs are powerful tools for modeling contextual information, as evidenced by their success at the task of language modeling. However, modeling contexts in very high dimensional space can lead to poor generalizability. We introduce the…

Computation and Language · Computer Science 2018-08-29 Sachin Mehta , Rik Koncel-Kedziorski , Mohammad Rastegari , Hannaneh Hajishirzi

Versatile and adaptive semantic understanding would enable autonomous systems to comprehend and interact with their surroundings. Existing fixed-class models limit the adaptability of indoor mobile and assistive autonomous systems. In this…

Robotics · Computer Science 2024-03-06 Christina Kassab , Matias Mattamala , Lintong Zhang , Maurice Fallon

Spatial cognition is essential for human intelligence, enabling problem-solving through visual simulations rather than solely relying on verbal reasoning. However, existing AI benchmarks primarily assess verbal reasoning, neglecting the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Linjie Li , Mahtab Bigverdi , Jiawei Gu , Zixian Ma , Yinuo Yang , Ziang Li , Yejin Choi , Ranjay Krishna

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

Sound · Computer Science 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for…

Computation and Language · Computer Science 2025-08-05 Wanqi Yang , Yanda Li , Yunchao Wei , Meng Fang , Ling Chen

This paper introduces the Learned User Significance Tracker (LUST), a framework designed to analyze video content and quantify the thematic relevance of its segments in relation to a user-provided textual description of significance. LUST…

Multimedia · Computer Science 2025-08-07 Anderson de Lima Luiz

Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intelligence of the future…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-15 Desh Raj

Speech separation approaches for single-channel, dry speech mixtures have significantly improved. However, real-world spatial and reverberant acoustic environments remain challenging, limiting the effectiveness of these approaches for…

Speech sound disorder is among the most common communication challenges in preschool children. Home-based practice is essential for effective therapy and for acquiring generalization of target sounds, yet sustaining engaging and consistent…

Human-Computer Interaction · Computer Science 2025-10-29 Sumin Hong , Xavier Briggs , Qingxiao Zheng , Yao Du , Jinjun Xiong , Toby Jia-jun Li

3D spatial reasoning in dynamic, audio-visual environments is a cornerstone of human cognition yet remains largely unexplored by existing Audio-Visual Large Language Models (AV-LLMs) and benchmarks, which predominantly focus on static or 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Mingfei Chen , Zijun Cui , Xiulong Liu , Jinlin Xiang , Caleb Zheng , Jingyuan Li , Eli Shlizerman

Multi-channel speech enhancement utilizes spatial information from multiple microphones to extract the target speech. However, most existing methods do not explicitly model spatial cues, instead relying on implicit learning from…

Sound · Computer Science 2023-09-20 Jiahui Pan , Shulin He , Hui Zhang , Xueliang Zhang
‹ Prev 1 8 9 10 Next ›