中文
相关论文

相关论文: MM-ALT: A Multimodal Automatic Lyric Transcription…

200 篇论文

Automatic Music Transcription (AMT) is a vital technology in the field of music information processing. Despite recent enhancements in performance due to machine learning techniques, current methods typically attain high accuracy in domains…

声音 · 计算机科学 2024-07-04 Gakusei Sato , Taketo Akama

Multi-instrument music transcription aims to convert polyphonic music recordings into musical scores assigned to each instrument. This task is challenging for modeling as it requires simultaneously identifying multiple instruments and…

音频与语音处理 · 电气工程与系统科学 2024-08-02 Sungkyun Chang , Emmanouil Benetos , Holger Kirchhoff , Simon Dixon

Integrating audio encoders with LLMs through connectors has enabled these models to process and comprehend audio modalities, significantly enhancing speech-to-text tasks, including automatic speech recognition (ASR) and automatic speech…

音频与语音处理 · 电气工程与系统科学 2024-09-18 Hongfei Xue , Wei Ren , Xuelong Geng , Kun Wei , Longhao Li , Qijie Shao , Linju Yang , Kai Diao , Lei Xie

Automatic Music Transcription (AMT), aiming to get musical notes from raw audio, typically uses frame-level systems with piano-roll outputs or language model (LM)-based systems with note-level predictions. However, frame-level systems…

声音 · 计算机科学 2025-01-08 Dichucheng Li , Yongyi Zang , Qiuqiang Kong

This paper presents MAST, a new model for Multimodal Abstractive Text Summarization that utilizes information from all three modalities -- text, audio and video -- in a multimodal video. Prior work on multimodal abstractive text…

计算与语言 · 计算机科学 2020-10-19 Aman Khullar , Udit Arora

Automatic music transcription (AMT) is the problem of analyzing an audio recording of a musical piece and detecting notes that are being played. AMT is a challenging problem, particularly when it comes to polyphonic music. The goal of AMT…

声音 · 计算机科学 2025-05-08 Yohannis Telila , Tommaso Cucinotta , Davide Bacciu

Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligned training pairs with precise frame-level annotations. Since…

声音 · 计算机科学 2025-11-19 Jonathan Yaffe , Ben Maman , Meinard Müller , Amit H. Bermano

Audio-to-score alignment (A2SA) is a multimodal task consisting in the alignment of audio signals to music scores. Recent literature confirms the benefits of Automatic Music Transcription (AMT) for A2SA at the frame-level. In this work, we…

声音 · 计算机科学 2022-01-03 Federico Simonetta , Stavros Ntalampiras , Federico Avanzini

Text-only adaptation of a transducer model remains challenging for end-to-end speech recognition since the transducer has no clearly separated acoustic model (AM), language model (LM) or blank model. In this work, we propose a modular…

Recent voice assistants are usually based on the cascade spoken language understanding (SLU) solution, which consists of an automatic speech recognition (ASR) engine and a natural language understanding (NLU) system. Because such approach…

计算与语言 · 计算机科学 2023-06-14 Anderson R. Avila , Mehdi Rezagholizadeh , Chao Xing

Speech-LLM models have demonstrated great performance in multi-modal and multi-task speech understanding. A typical speech-LLM paradigm is integrating speech modality with a large language model (LLM). While the Whisper encoder was…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Wei Liu , Jiahong Li , Yiwen Shao , Dong Yu

Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal…

声音 · 计算机科学 2024-06-04 Josh Gardner , Simon Durand , Daniel Stoller , Rachel M. Bittner

We propose a cloud-based multimodal dialog platform for the remote assessment and monitoring of Amyotrophic Lateral Sclerosis (ALS) at scale. This paper presents our vision, technology setup, and an initial investigation of the efficacy of…

Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a unified…

计算与语言 · 计算机科学 2021-09-15 Renjie Zheng , Junkun Chen , Mingbo Ma , Liang Huang

Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific…

Most of the current supervised automatic music transcription (AMT) models lack the ability to generalize. This means that they have trouble transcribing real-world music recordings from diverse musical genres that are not presented in the…

声音 · 计算机科学 2021-07-30 Kin Wai Cheuk , Dorien Herremans , Li Su

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU-LLaMA), capable of…

声音 · 计算机科学 2023-08-23 Shansong Liu , Atin Sakkeer Hussain , Chenshuo Sun , Ying Shan

Auscultation is a vital diagnostic tool, yet its utility is often limited by subjective interpretation. While general-purpose Audio-Language Models (ALMs) excel in general domains, they struggle with the nuances of physiological signals. We…

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder…

人工智能 · 计算机科学 2023-12-15 Haoyu Zhang , Yu Wang , Guanghao Yin , Kejun Liu , Yuanyuan Liu , Tianshu Yu

This paper presents Audio-Visual LLM, a Multimodal Large Language Model that takes both visual and auditory inputs for holistic video understanding. A key design is the modality-augmented training, which involves the integration of…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Fangxun Shu , Lei Zhang , Hao Jiang , Cihang Xie