中文
相关论文

相关论文: AudioTime: A Temporally-aligned Audio-text Benchma…

200 篇论文

Recently, audio generation tasks have attracted considerable research interests. Precise temporal controllability is essential to integrate audio generation with real applications. In this work, we propose a temporal controlled audio…

声音 · 计算机科学 2024-07-18 Zeyu Xie , Xuenan Xu , Zhizheng Wu , Mengyue Wu

Recent advancements in machine learning have fueled research on multimodal tasks, such as for instance text-to-video and text-to-audio retrieval. These tasks require models to understand the semantic content of video and audio data,…

信息检索 · 计算机科学 2024-09-04 Andreea-Maria Oncescu , João F. Henriques , A. Sophia Koepke

Automated audio captioning aims at generating natural language descriptions for given audio clips, not only detecting and classifying sounds, but also summarizing the relationships between audio events. Recent research advances in audio…

声音 · 计算机科学 2024-07-19 Zeyu Xie , Xuenan Xu , Mengyue Wu , Kai Yu

Recent advances in generative video models have enabled the creation of high-quality videos based on natural language prompts. However, these models frequently lack fine-grained temporal control, meaning they do not allow users to specify…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Shira Schiber , Ofir Lindenbaum , Idan Schwartz

While recent work in controllable text-to-audio (TTA) generation has achieved fine-grained control through timestamp conditioning, its scope remains limited by audio quality and input format. These models often suffer from poor audio…

声音 · 计算机科学 2025-10-14 Zihao Zheng , Zeyu Xie , Xuenan Xu , Wen Wu , Chao Zhang , Mengyue Wu

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the…

声音 · 计算机科学 2023-12-29 Zhifang Guo , Jianguo Mao , Rui Tao , Long Yan , Kazushige Ouchi , Hong Liu , Xiangdong Wang

Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text-conditioned audio generation. Existing contrastive…

音频与语音处理 · 电气工程与系统科学 2025-05-13 Paul Primus , Florian Schmid , Gerhard Widmer

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g.,…

声音 · 计算机科学 2025-12-15 Hualei Wang , Yiming Li , Shuo Ma , Hong Liu , Xiangdong Wang

Large Audio Language Models (LALMs) are increasingly applied to audio understanding and multimodal reasoning, yet their ability to locate when events occur remains underexplored. We present the first systematic study of temporal bias in…

计算与语言 · 计算机科学 2025-10-15 Jiayu Yao , Shenghua Liu , Yiwei Wang , Rundong Cheng , Lingrui Mei , Baolong Bi , Zhen Xiong , Xueqi Cheng

Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such…

声音 · 计算机科学 2024-03-08 Xuenan Xu , Xiaohang Xu , Zeyu Xie , Pingyue Zhang , Mengyue Wu , Kai Yu

While Large Audio Language Models (LALMs) achieve strong performance on short audio, they degrade on long-form inputs. This degradation is more severe in temporal awareness tasks, where temporal alignment becomes increasingly inaccurate as…

音频与语音处理 · 电气工程与系统科学 2026-04-27 Mingchen Shao , Hang Su , Wenjie Tian , Bingshen Mu , Zhennan Lin , Lichun Fan , Zhenbo Luo , Jian Luan , Lei Xie

Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consistency due to limited natural language understanding and data…

Multi-modal contrastive learning techniques in the audio-text domain have quickly become a highly active area of research. Most works are evaluated with standard audio retrieval and classification benchmarks assuming that (i) these models…

声音 · 计算机科学 2023-03-21 Ho-Hsiang Wu , Oriol Nieto , Juan Pablo Bello , Justin Salamon

Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale…

声音 · 计算机科学 2026-04-21 Yuxuan Jiang , Zehua Chen , Zeqian Ju , Yusheng Dai , Weibei Dou , Jun Zhu

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

声音 · 计算机科学 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang

Video large language models have achieved remarkable performance in tasks such as video question answering, however, their temporal understanding remains suboptimal. To address this limitation, we curate a dedicated instruction fine-tuning…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Yunxiao Wang , Meng Liu , Wenqi Liu , Xuemeng Song , Bin Wen , Fan Yang , Tingting Gao , Di Zhang , Guorui Zhou , Liqiang Nie

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control…

声音 · 计算机科学 2026-02-10 Yisu Liu , Chenxing Li , Wanqian Zhang , Wenfu Wang , Meng Yu , Ruibo Fu , Zheng Lin , Weiping Wang , Dong Yu

Conversational Spoken Language Models (SLMs) are emerging as a promising paradigm for real-time speech interaction. However, their capacity of temporal dynamics, including the ability to manage timing, tempo and simultaneous speaking,…

音频与语音处理 · 电气工程与系统科学 2026-05-04 Kai-Wei Chang , En-Pei Hu , Chun-Yi Kuan , Wenze Ren , Wei-Chih Chen , Guan-Ting Lin , Yu Tsao , Shao-Hua Sun , Hung-yi Lee , James Glass

Recent advances in text-to-audio (TTA) generation excel at synthesizing short audio clips but struggle with long-form narrative audio, which requires temporal coherence and compositional reasoning. To address this gap, we propose…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Yuxin Guo , Teng Wang , Yuying Ge , Shijie Ma , Yixiao Ge , Wei Zou , Ying Shan

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image editing, audio-driven…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Kaixin Shen , Ruijie Quan , Linchao Zhu , Jun Xiao , Yi Yang
‹ 上一页 1 2 3 10 下一页 ›