中文
相关论文

相关论文: AudioTime: A Temporally-aligned Audio-text Benchma…

200 篇论文

Multimodal question answering tasks can be used as proxy tasks to study systems that can perceive and reason about the world. Answering questions about different types of input modalities stresses different aspects of reasoning such as…

计算与语言 · 计算机科学 2019-11-22 Haytham M. Fayek , Justin Johnson

Large Audio-Language Models (LALMs) perform well on audio understanding tasks but lack multistep reasoning and tool-calling found in recent Large Language Models (LLMs). This paper presents AudioToolAgent, a framework that coordinates…

声音 · 计算机科学 2026-02-16 Gijs Wijngaard , Elia Formisano , Michel Dumontier , Jenia Jitsev

The advancement of Machine learning (ML), Large Audio Language Models (LALMs), and autonomous AI agents in Music Information Retrieval (MIR) necessitates a shift from static tagging to rich, human-aligned representation learning. However,…

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval,…

声音 · 计算机科学 2026-01-30 Hui Wang , Jinghua Zhao , Junyang Cheng , Cheng Liu , Yuhang Jia , Haoqin Sun , Jiaming Zhou , Yong Qin

Cross-modal (e.g. image-text, video-text) retrieval is an important task in information retrieval and multimodal vision-language understanding field. Temporal understanding makes video-text retrieval more challenging than image-text…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yang Du , Yuqi Liu , Qin Jin

Most modern approaches for audio processing are opaque, in the sense that they do not provide an explanation for their decisions. For this reason, various methods have been proposed to explain the outputs generated by these models. Good…

声音 · 计算机科学 2025-10-21 Cecilia Bolaños , Leonardo Pepino , Martin Meza , Luciana Ferrer

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear…

The recent rapid advancements in language models (LMs) have garnered attention in medical time series-text multimodal learning. However, existing contrastive learning-based and prompt-based LM approaches tend to be biased, often assigning a…

机器学习 · 计算机科学 2025-09-09 Jiexia Ye , Weiqi Zhang , Ziyue Li , Jia Li , Meng Zhao , Fugee Tsung

State-of-the-art models of lexical semantic change detection suffer from noise stemming from vector space alignment. We have empirically tested the Temporal Referencing method for lexical semantic change and show that, by avoiding…

计算与语言 · 计算机科学 2020-07-23 Haim Dubossarsky , Simon Hengchen , Nina Tahmasebi , Dominik Schlechtweg

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

声音 · 计算机科学 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models…

声音 · 计算机科学 2026-04-28 Yi Yuan , Xubo Liu , Haohe Liu , Xiyuan Kang , Zhuo Chen , Yuxuan Wang , Mark D. Plumbley , Wenwu Wang

Large Language Models (LLMs) have made significant strides in text generation and comprehension, with recent advancements extending into multimodal LLMs that integrate visual and audio inputs. However, these models continue to struggle with…

计算与语言 · 计算机科学 2024-10-17 Arushi Goel , Karan Sapra , Matthieu Le , Rafael Valle , Andrew Tao , Bryan Catanzaro

Text-to-audio generation (TTA) produces audio from a text description, learning from pairs of audio samples and hand-annotated text. However, commercializing audio generation is challenging as user-input prompts are often under-specified…

Developing tools to automatically detect check-worthy claims in political debates and speeches can greatly help moderators of debates, journalists, and fact-checkers. While previous work on this problem has focused exclusively on the text…

计算与语言 · 计算机科学 2024-01-19 Petar Ivanov , Ivan Koychev , Momchil Hardalov , Preslav Nakov

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

声音 · 计算机科学 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the…

音频与语音处理 · 电气工程与系统科学 2026-03-26 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Junbo Zhang , Jian Luan

In this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in…

音频与语音处理 · 电气工程与系统科学 2025-08-05 Hokuto Munakata , Taichi Nishimura , Shota Nakada , Tatsuya Komatsu

Human experts typically integrate numerical and textual multimodal information to analyze time series. However, most traditional deep learning predictors rely solely on unimodal numerical data, using a fixed-length window for training and…

计算与语言 · 计算机科学 2024-12-17 Chengsen Wang , Qi Qi , Jingyu Wang , Haifeng Sun , Zirui Zhuang , Jinming Wu , Lei Zhang , Jianxin Liao

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Weixia Zhang , Chengguang Zhu , Jingnan Gao , Yichao Yan , Guangtao Zhai , Xiaokang Yang

We propose an efficient workflow for high-quality offline alignment of in-the-wild performance audio and corresponding sheet music scans (images). Recent work on audio-to-score alignment extends dynamic time warping (DTW) to be…

声音 · 计算机科学 2024-11-13 Irmak Bukey , Michael Feffer , Chris Donahue