中文
相关论文

相关论文: The Sonar Moment: Benchmarking Audio-Language Mode…

200 篇论文

Large Audio Language Models (LALMs) have made significant progress. While increasingly deployed in real-world applications, LALMs face growing safety risks from jailbreak attacks that bypass safety alignment. However, there remains a lack…

密码学与安全 · 计算机科学 2026-03-03 Zifan Peng , Yule Liu , Zhen Sun , Mingchen Li , Zeren Luo , Jingyi Zheng , Wenhan Dong , Xinlei He , Xuechao Wang , Yingjie Xue , Shengmin Xu , Xinyi Huang

We introduce Voices of Civilizations, the first multilingual QA benchmark for evaluating audio LLMs' cultural comprehension on full-length music recordings. Covering 380 tracks across 38 languages, our automated pipeline yields 1,190…

声音 · 计算机科学 2026-03-03 Shangda Wu , Ziya Zhou , Yongyi Zang , Yutong Zheng , Dafang Liang , Ruibin Yuan , Qiuqiang Kong

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the…

Autoregressive (AR) modeling is invaluable in signal processing, in particular in speech and audio fields. Attempts in the literature can be found that regularize or constrain either the time-domain signal values or the AR coefficients,…

音频与语音处理 · 电气工程与系统科学 2026-02-06 Ondřej Mokrý , Pavel Rajmic

Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are…

Audio generation has achieved remarkable progress with the advance of sophisticated generative models, such as diffusion models (DMs) and autoregressive (AR) models. However, due to the naturally significant sequence length of audio, the…

声音 · 计算机科学 2024-12-18 Kai Qiu , Xiang Li , Hao Chen , Jie Sun , Jinglu Wang , Zhe Lin , Marios Savvides , Bhiksha Raj

The rise of Large Audio Language Models (LAMs) brings both potential and risks, as their audio outputs may contain harmful or unethical content. However, current research lacks a systematic, quantitative evaluation of LAM safety especially…

Recent Large Audio-Language Models (LALMs) have shown strong performance on various audio understanding tasks such as speech translation and Audio Q\&A. However, they exhibit significant limitations on challenging audio reasoning tasks in…

计算与语言 · 计算机科学 2025-09-29 Zhen Xiong , Yujun Cai , Zhecheng Li , Junsong Yuan , Yiwei Wang

Large Audio Language Models (LALMs) expand jailbreak risks from token-level prompting to the full speech perception-to-reasoning pipeline, where unsafe behavior can be induced through semantics, acoustic style, signal artifacts, or internal…

声音 · 计算机科学 2026-05-29 Bo-Han Feng , Yu-Hsuan Li Liang , Chien-Feng Liu , You-Hsuan Chang , Yun-Nung Chen

The emergence of Vision-Language Models (VLMs) has introduced new paradigms for global image geo-localization through retrieval-augmented generation (RAG) and reasoning-driven inference. However, RAG methods are constrained by retrieval…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Bo Yu , Fengze Yang , Yiming Liu , Chao Wang , Xuewen Luo , Taozhe Li , Ruimin Ke , Xiaofan Zhou , Chenxi Liu

Sound recognition is an important and popular function of smart devices. The location of sound is basic information associated with the acoustic source. Apart from sound recognition, whether the acoustic sources can be localized largely…

声音 · 计算机科学 2022-10-03 Weiguo Wang , Jinming Li , Yuan He , Yunhao Liu

Geographic reasoning is a fundamental cognitive capability that requires models to infer plausible locations by synthesizing visual evidence with spatial world knowledge. Despite recent advances in large vision-language models (LVLMs),…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Pengyue Jia , Yingyi Zhang , Xiangyu Zhao , Sharon Li

Tool-augmented Large Language Models (LLMs) have shown impressive capabilities in remote sensing (RS) applications. However, existing benchmarks assume question-answering input templates over predefined image-text data pairs. These…

计算与语言 · 计算机科学 2024-05-03 Simranjit Singh , Michael Fore , Dimitrios Stamoulis

Vision-language models (VLMs) have demonstrated strong performance in image geolocation, a capability further sharpened by frontier multimodal large reasoning models (MLRMs). This poses a significant privacy risk, as these widely accessible…

密码学与安全 · 计算机科学 2026-02-19 Ruixin Yang , Ethan Mendes , Arthur Wang , James Hays , Sauvik Das , Wei Xu , Alan Ritter

Vision-language models (VLMs) have advanced rapidly, yet their capacity for image-grounded geolocation in open-world conditions, a task that is challenging and of demand in real life, has not been comprehensively evaluated. We present…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Zhaofang Qian , Hardy Chen , Zeyu Wang , Li Zhang , Zijun Wang , Xiaoke Huang , Hui Liu , Xianfeng Tang , Zeyu Zheng , Haoqin Tu , Cihang Xie , Yuyin Zhou

The next Point-of-Interest (POI) recommendation task aims to predict users' next destinations based on their historical movement data and plays a key role in location-based services and personalized applications. Accurate next POI…

信息检索 · 计算机科学 2025-05-21 Zhao Liu , Wei Liu , Huajie Zhu , Jianxing Yu , Jian Yin , Wang-Chien Lee , Shun Wang

Large Audio Language Models (LALM) combine the audio perception models and the Large Language Models (LLM) and show a remarkable ability to reason about the input audio, infer the meaning, and understand the intent. However, these systems…

音频与语音处理 · 电气工程与系统科学 2024-11-26 Saurabhchand Bhati , Yuan Gong , Leonid Karlinsky , Hilde Kuehne , Rogerio Feris , James Glass

With a fast developing pace of geographic applications, automatable and intelligent models are essential to be designed to handle the large volume of information. However, few researchers focus on geographic natural language processing, and…

计算与语言 · 计算机科学 2023-05-12 Dongyang Li , Ruixue Ding , Qiang Zhang , Zheng Li , Boli Chen , Pengjun Xie , Yao Xu , Xin Li , Ning Guo , Fei Huang , Xiaofeng He

Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often seek high audio-visual similarities for likely positive…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Shentong Mo , Pedro Morgado

Large Audio Language Models (LALMs) have significantly advanced audio understanding but introduce critical security risks, particularly through audio jailbreaks. While prior work has focused on English-centric attacks, we expose a far more…

声音 · 计算机科学 2025-04-03 Jaechul Roh , Virat Shejwalkar , Amir Houmansadr