中文

Story2MIDI:基于文本生成情感对齐的音乐

声音 2025-12-03 v1 人工智能 计算与语言

摘要

本文提出 Story2MIDI,一个基于序列到序列 Transformer 的模型,用于从给定文本生成情感对齐的音乐。为开发该模型,我们通过合并情感分析文本数据集和音乐情感分类数据集构建了 Story2MIDI 数据集。 resulting dataset contains pairs of text blurbs and music pieces that evoke the same emotions in the reader or listener. Despite the small scale of our dataset and limited computational resources, our results indicate that our model effectively learns emotion-relevant features in music and incorporates them into its generation process, producing samples with diverse emotional responses. 我们使用客观音乐指标和人类听觉研究进行评估,确认了模型捕获预期情感线索的能力。

关键词

引用

@article{arxiv.2512.02192,
  title  = {Story2MIDI: Emotionally Aligned Music Generation from Text},
  author = {Mohammad Shokri and Alexandra C. Salem and Gabriel Levine and Johanna Devaney and Sarah Ita Levitan},
  journal= {arXiv preprint arXiv:2512.02192},
  year   = {2025}
}

备注

8 pages (6 pages of main text + 2 pages of references and appendices), 4 figures, 1 table. Presented at IEEE Big Data 2025 3rd Workshop on AI Music Generation (AIMG 2025)