中文
相关论文

相关论文: Eureka-Audio: Triggering Audio Intelligence in Com…

200 篇论文

Recent advances in large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged as a promising alternative to the autoregressive paradigm,…

Benchmarking plays a pivotal role in assessing and enhancing the performance of compact deep learning models designed for execution on resource-constrained devices, such as microcontrollers. Our study introduces a novel, entirely…

声音 · 计算机科学 2024-03-18 René Groh , Nina Goes , Andreas M. Kist

This survey overviews various meta-learning approaches used in audio and speech processing scenarios. Meta-learning is used where model performance needs to be maximized with minimum annotated samples, making it suitable for low-sample…

声音 · 计算机科学 2025-03-14 Athul Raimon , Shubha Masti , Shyam K Sateesh , Siyani Vengatagiri , Bhaskarjyoti Das

Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate LALMs' audio understanding performance, researchers…

声音 · 计算机科学 2026-02-16 Han Yin , Jung-Woo Choi

Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-based free-form…

Audio large language models (Audio LLMs) demonstrate strong performance on speech understanding tasks, yet their ability to understand paralinguistic information remains limited. To systematically quantify this issue, we introduce…

声音 · 计算机科学 2026-05-28 Jiacheng Pang , Ashutosh Chaubey , Mohammad Soleymani

End-to-end models have achieved significant improvement on automatic speech recognition. One common method to improve performance of these models is expanding the data-space through data augmentation. Meanwhile, human auditory inspired…

音频与语音处理 · 电气工程与系统科学 2022-04-12 Zehai Tu , Jack Deadman , Ning Ma , Jon Barker

Recent years have witnessed significant progress in multilingual automatic speech recognition (ASR), driven by the emergence of end-to-end (E2E) models and the scaling of multilingual datasets. Despite that, two main challenges persist in…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Zheshu Song , Jianheng Zhuo , Yifan Yang , Ziyang Ma , Shixiong Zhang , Xie Chen

Speech enhancement is crucial for ubiquitous human-computer interaction. Recently, ultrasound-based acoustic sensing has emerged as an attractive choice for speech enhancement because of its superior ubiquity and performance. However, due…

声音 · 计算机科学 2025-05-20 Luca Jiang-Tao Yu , Running Zhao , Sijie Ji , Edith C. H. Ngai , Chenshu Wu

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

计算与语言 · 计算机科学 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Aligning pretrained audio encoders and Large Language Models (LLMs) offers a promising, parameter-efficient path to building powerful multimodal agents. However, existing methods often require costly full-model finetuning or rely on static…

声音 · 计算机科学 2025-10-16 Ruitao Feng , Bixi Zhang , Sheng Liang , Zheng Yuan

A universal audio representation should capture fine-grained speech cues and high-level semantics for environmental sounds and music in a single encoder. Existing encoders often excel in one domain but degrade in others. We propose…

声音 · 计算机科学 2026-03-10 Yuxuan Chen , Peize He , Haoyuan Yu , Junzi Zhang

Large Audio Language Models (LALMs) have demonstrated strong capabilities in audio understanding and reasoning. However, their performance on fine grained auditory perception remains unreliable, and existing approaches largely rely on data…

声音 · 计算机科学 2026-02-12 Liyang Chen , Hongkai Chen , Yujun Cai , Sifan Li , Qingwen Ye , Yiwei Wang

Self-Supervised Learning (SSL) models have demonstrated exceptional performance in various speech tasks, particularly in low-resource and multilingual domains. Recent works show that fusing diverse SSL models could achieve superior…

声音 · 计算机科学 2024-06-07 Tejes Srivastava , Jiatong Shi , William Chen , Shinji Watanabe

We propose a multi-stage framework for universal speech enhancement, designed for the Interspeech 2025 URGENT Challenge. Our system first employs a Sparse Compression Network to robustly separate sources and extract an initial clean speech…

声音 · 计算机科学 2025-06-03 Nabarun Goswami , Tatsuya Harada

Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPASE, an extension of the low-hallucination PASE framework tailored for USE. At its core is…

音频与语音处理 · 电气工程与系统科学 2026-04-17 Xiaobin Rong , Zheng Wang , Yushi Wang , Jun Gao , Jing Lu

We present Voice Evaluation of Reasoning Ability (VERA), a benchmark for evaluating reasoning ability in voice-interactive systems under real-time conversational constraints. VERA comprises 2,931 voice-native episodes derived from…

音频与语音处理 · 电气工程与系统科学 2025-10-01 Yueqian Lin , Zhengmian Hu , Qinsi Wang , Yudong Liu , Hengfan Zhang , Jayakumar Subramanian , Nikos Vlassis , Hai Helen Li , Yiran Chen

Generative modeling has recently achieved remarkable success across text, image, and audio domains, demonstrating powerful capabilities for unified representation learning. However, audio generation models still face challenges in terms of…

声音 · 计算机科学 2025-10-31 Chengwei Liu , Haoyin Yan , Shaofei Xue , Xiaotao Liang , Yinghao Liu , Zheng Xue , Gang Song , Boyang Zhou

Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fails for reasoning LLMs (RLMs) whose built-in…

计算与语言 · 计算机科学 2026-03-11 Petr Grinberg , Hassan Shahmohammadi

Self-supervised language and audio models effectively predict brain responses to speech. However, traditional prediction models rely on linear mappings from unimodal features, despite the complex integration of auditory signals with…

计算与语言 · 计算机科学 2025-02-19 Danny Dongyeop Han , Yunju Cho , Jiook Cha , Jay-Yoon Lee