中文
相关论文

相关论文: Towards Spatial Audio Understanding via Question A…

200 篇论文

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering…

音频与语音处理 · 电气工程与系统科学 2023-12-19 Yuzhu Wang , Archontis Politis , Tuomas Virtanen

We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED), audio classification, and direction-of-arrival estimation…

音频与语音处理 · 电气工程与系统科学 2026-04-16 Dongheon Lee , Younghoo Kwon , Jung-Woo Choi

Detecting auditory attention based on brain signals enables many everyday applications, and serves as part of the solution to the cocktail party effect in speech processing. Several studies leverage the correlation between brain signals and…

人机交互 · 计算机科学 2024-10-28 Siqi Cai , Pengcheng Sun , Tanja Schultz , Haizhou Li

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhibit significant…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Kexin Tian , Jingrui Mao , Yunlong Zhang , Jiwan Jiang , Yang Zhou , Zhengzhong Tu

This work introduces a feature extracted from stereophonic/binaural audio signals aiming to represent a measure of perceived quality degradation in processed spatial auditory scenes. The feature extraction technique is based on a simplified…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Pablo M. Delgado , Jürgen Herre

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Most modern approaches for audio processing are opaque, in the sense that they do not provide an explanation for their decisions. For this reason, various methods have been proposed to explain the outputs generated by these models. Good…

声音 · 计算机科学 2025-10-21 Cecilia Bolaños , Leonardo Pepino , Martin Meza , Luciana Ferrer

Contrastive language--audio pretraining (CLAP) has achieved remarkable success as an audio--text embedding framework, but existing approaches are limited to monaural or single-source conditions and cannot fully capture spatial information.…

3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existing research has primarily focused on indoor household tasks…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Penglei Sun , Yaoxian Song , Xiang Liu , Xiaofei Yang , Qiang Wang , Tiefeng Li , Yang Yang , Xiaowen Chu

This paper investigates the effectiveness of factorial speech processing models in noise-robust automatic speech recognition tasks. For this purpose, the paper proposes an idealistic approach for modeling state-conditional observation…

机器学习 · 计算机科学 2016-10-06 Mahdi Khademian , Mohammad Mehdi Homayounpour

Spatial reasoning plays a vital role in both human cognition and machine intelligence, prompting new research into language models' (LMs) capabilities in this regard. However, existing benchmarks reveal shortcomings in evaluating…

计算与语言 · 计算机科学 2024-05-27 Fangjun Li , David C. Hogg , Anthony G. Cohn

Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identify this error pattern as the evidence bottleneck:…

声音 · 计算机科学 2026-05-29 Xinyuan Xie , Shunian Chen , Zhiheng Liu , Yuhao Zhang , Zhiqiang Lv , Liyin Liang , Benyou Wang

Geospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches. While existing geospatial QA datasets exist, they are limited in both scale and diversity, often relying solely on textual…

计算与语言 · 计算机科学 2025-03-12 Zekun Li , Malcolm Grossman , Eric , Qasemi , Mihir Kulkarni , Muhao Chen , Yao-Yi Chiang

When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be…

多媒体 · 计算机科学 2020-08-19 Ying Cheng , Ruize Wang , Zhihao Pan , Rui Feng , Yuejie Zhang

Recently, Large Language Models (LLMs) have introduced a novel paradigm in Time Series Analysis (TSA), leveraging strong language capabilities to support tasks such as forecasting and anomaly detection. However, these analysis tasks cannot…

机器学习 · 计算机科学 2026-05-11 Wei Li , Zhe Xie , Yuxuan Liang , Xinli Hao , Yunyao Cheng , Dan Pei , Xiaofeng Meng

We present a Chain-of-Action (CoA) framework for multimodal and retrieval-augmented Question-Answering (QA). Compared to the literature, CoA overcomes two major challenges of current QA applications: (i) unfaithful hallucination that is…

计算与语言 · 计算机科学 2025-02-24 Zhenyu Pan , Haozheng Luo , Manling Li , Han Liu

Recent advances in object-centric representation learning have shown that slot attention-based methods can effectively decompose visual scenes into object slot representations without supervision. However, existing approaches typically…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Huankun Sheng , Ming Li , Yixiang Wei , Yeying Fan , Yu-Hui Wen , Tieliang Gong , Yong-Jin Liu

Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently…

声音 · 计算机科学 2026-02-27 Townim Faisal Chowdhury , Ta Duc Huy , Siqi Pan , Jeremy Stoddard , Zhibin Liao

Visual Question and Answering (VQA) problems are attracting increasing interest from multiple research disciplines. Solving VQA problems requires techniques from both computer vision for understanding the visual contents of a presented…

计算机视觉与模式识别 · 计算机科学 2016-04-07 Ilija Ilievski , Shuicheng Yan , Jiashi Feng

Supervised learning methods have shown effectiveness in estimating spatial acoustic parameters such as time difference of arrival, direct-to-reverberant ratio and reverberation time. However, they still suffer from the simulation-to-reality…

声音 · 计算机科学 2024-09-10 Bing Yang , Xiaofei Li