中文
相关论文

相关论文: LLM-AD: Large Language Model based Audio Descripti…

200 篇论文

Automatic music captioning, which generates natural language descriptions for given music tracks, holds significant potential for enhancing the understanding and organization of large volumes of musical data. Despite its importance,…

声音 · 计算机科学 2023-08-01 SeungHeon Doh , Keunwoo Choi , Jongpil Lee , Juhan Nam

In this work, we propose the use of "aligned visual captions" as a mechanism for integrating information contained within videos into retrieval augmented generation (RAG) based chat assistant systems. These captions are able to describe the…

人工智能 · 计算机科学 2024-05-29 Kevin Dela Rosa

Audiobook generation aims to create rich, immersive listening experiences from multimodal inputs, but current approaches face three critical challenges: (1) the lack of synergistic generation of diverse audio types (e.g., speech, sound…

声音 · 计算机科学 2025-08-13 Yan Rong , Shan Yang , Chenxing Li , Dong Yu , Li Liu

Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as images or audio.…

计算与语言 · 计算机科学 2025-11-13 Yiming Gao , Bin Wang , Chengwei Wei , Shuo Sun , AiTi Aw

Large Language Models (LLMs) have become a crucial tool in Visual Question Answering (VQA) for handling knowledge-intensive questions in few-shot or zero-shot scenarios. However, their reliance on massive training datasets often causes them…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Quanxing Xu , Ling Zhou , Feifei Zhang , Jinyu Tian , Rubing Huang

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

多媒体 · 计算机科学 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Recent advancements in Large Language Models (LLMs) offer new opportunities to create natural language interfaces for Autonomous Driving Systems (ADSs), moving beyond rigid inputs. This paper addresses the challenge of mapping the…

机器人学 · 计算机科学 2026-01-26 Marvin Seegert , Korbinian Moller , Johannes Betz

Virtual Reality (VR) has emerged as a powerful tool for workforce training, offering immersive, interactive, and risk-free environments that enhance skill acquisition, decision-making, and confidence. Despite its advantages, developing VR…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Subin Raj Peter

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice ear-piece…

计算与语言 · 计算机科学 2016-12-01 Hans Krupakar , Keerthika Rajvel , Bharathi B , Angel Deborah S , Vallidevi Krishnamurthy

Audio deepfake detection (ADD) has grown increasingly important due to the rise of high-fidelity audio generative models and their potential for misuse. Given that audio large language models (ALLMs) have made significant progress in…

声音 · 计算机科学 2025-07-09 Hao Gu , Jiangyan Yi , Chenglong Wang , Jianhua Tao , Zheng Lian , Jiayi He , Yong Ren , Yujie Chen , Zhengqi Wen

We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation…

The design flow of processors, particularly in hardware description languages (HDL) like Verilog and Chisel, is complex and costly. While recent advances in large language models (LLMs) have significantly improved coding tasks in software…

The goal of audio captioning is to translate input audio into its description using natural language. One of the problems in audio captioning is the lack of training data due to the difficulty in collecting audio-caption pairs by crawling…

音频与语音处理 · 电气工程与系统科学 2020-12-15 Yuma Koizumi , Yasunori Ohishi , Daisuke Niizumi , Daiki Takeuchi , Masahiro Yasuda

Automatic differentiation (AD) is a technique for computing the derivative of a function represented by a program. This technique is considered as the de-facto standard for computing the differentiation in many machine learning and…

Multimodal Large Language Models (MLLMs) mimic human perception and reasoning system by integrating powerful Large Language Models (LLMs) with various modality encoders (e.g., vision, audio), positioning LLMs as the "brain" and various…

计算机视觉与模式识别 · 计算机科学 2024-08-29 Jiaxing Huang , Jingyi Zhang

This survey and application guide to multimodal large language models(MLLMs) explores the rapidly developing field of MLLMs, examining their architectures, applications, and impact on AI and Generative Models. Starting with foundational…

人工智能 · 计算机科学 2025-12-02 Chia Xin Liang , Pu Tian , Caitlyn Heqi Yin , Yao Yua , Wei An-Hou , Li Ming , Xinyuan Song , Tianyang Wang , Ziqian Bi , Ming Liu

Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning and open-world generalization. However, the excessive computational overhead and high…

机器人学 · 计算机科学 2026-05-26 Ruoyu Yao , Ruiguo Zhong , Pei Liu , Mingxing Peng , Rui Yang , Jun Ma

Enforcing archival standards requires specialized expertise, and manually creating metadata descriptions for archival materials is a tedious and error-prone task. This work aims at exploring the potential of agentic AI and large language…

人工智能 · 计算机科学 2026-02-03 Jinghua Groppe , Andreas Marquet , Annabel Walz , Sven Groppe

Accurate dialogue description in audiovisual video captioning is crucial for downstream understanding and generation tasks. However, existing models generally struggle to produce faithful dialogue descriptions within audiovisual captions.…