中文
相关论文

相关论文: Finding Tori: Self-supervised Learning for Analyzi…

200 篇论文

Pansori is one of the most representative vocal genres of Korean traditional music, which has an elaborated vocal melody line with strong vibrato. Although the music is transmitted orally without any music notation, transcribing pansori…

声音 · 计算机科学 2024-10-18 Sangheon Park , Danbinaerin Han , Dasaem Jeong

Computational analysis of folk song audio is challenging due to structural irregularities and the need for manual annotation. We propose a method for automatic motive segmentation in Korean folk songs by fine-tuning a speech transcription…

声音 · 计算机科学 2025-08-15 Danbinaerin Han , Dasaem Jeong , Juhan Nam

In this paper, we focus on singing techniques within the scope of music information retrieval research. We investigate how singers use singing techniques using real-world recordings of famous solo singers in Japanese popular music songs…

声音 · 计算机科学 2022-11-17 Yuya Yamamoto , Juhan Nam , Hiroko Terasawa

Lyric translation, a field studied for over a century, is now attracting computational linguistics researchers. We identified two limitations in previous studies. Firstly, lyric translation studies have predominantly focused on Western…

计算与语言 · 计算机科学 2024-05-21 Haven Kim , Jongmin Jung , Dasaem Jeong , Juhan Nam

Common AI music composition algorithms based on artificial neural networks are to train a machine by feeding a large number of music pieces and create artificial neural networks that can produce music similar to the input music data. This…

声音 · 计算机科学 2022-03-30 Mai Lan Tran , Dongjin Lee , Jae-Hun Jung

Studying under-represented music traditions under the MIR scope is crucial, not only for developing novel analysis tools, but also for unveiling musical functions that might prove useful in studying world musics. This paper presents a…

Jeongganbo is a unique music representation invented by Sejong the Great. Contrary to the western music notation, the pitch of each note is encrypted and the length is visualized directly in a matrix form in Jeongganbo. We use topological…

声音 · 计算机科学 2021-07-02 Mai Lan Tran , Changbom Park , Jae-Hun Jung

Existing question answering systems mainly focus on dealing with text data. However, much of the data produced daily is stored in the form of tables that can be found in documents and relational databases, or on the web. To solve the task…

计算与语言 · 计算机科学 2022-05-03 Changwook Jun , Jooyoung Choi , Myoseop Sim , Hyun Kim , Hansol Jang , Kyungkoo Min

This paper presents a computational methodology for analyzing intonation and deriving tuning systems in microtonal oral traditions, utilizing pitch histograms, Dynamic Time Warping (DTW), and optimization techniques, with a case study on a…

声音 · 计算机科学 2025-08-29 Sepideh Shafiei , Shapour Hakam

The goal of this paper is twofold. First, we introduce DALI, a large and rich multimodal dataset containing 5358 audio tracks with their time-aligned vocal melody notes and lyrics at four levels of granularity. The second goal is to explain…

音频与语音处理 · 电气工程与系统科学 2019-06-26 Gabriel Meseguer-Brocal , Alice Cohen-Hadria , Geoffroy Peeters

Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient…

声音 · 计算机科学 2026-03-02 Jiajia Li , Jiliang Hu , Ziyi Pan , Chong Chen , Zuchao Li , Ping Wang , Lefei Zhang

We introduce a data-driven approach to automatic pitch correction of solo singing performances. The proposed approach predicts note-wise pitch shifts from the relationship between the respective spectrograms of the singing and…

声音 · 计算机科学 2020-02-25 Sanna Wager , George Tzanetakis , Cheng-i Wang , Minje Kim

Instruction tuning has emerged as a powerful technique, significantly boosting zero-shot performance on unseen tasks. While recent work has explored cross-lingual generalization by applying instruction tuning to multilingual models,…

计算与语言 · 计算机科学 2024-06-14 Janghoon Han , Changho Lee , Joongbo Shin , Stanley Jungkyu Choi , Honglak Lee , Kynghoon Bae

This paper presents a new large-scale Japanese speech corpus for training automatic speech recognition (ASR) systems. This corpus contains over 2,000 hours of speech with transcripts built on Japanese TV recordings and their subtitles. We…

声音 · 计算机科学 2021-03-30 Shintaro Ando , Hiromasa Fujihara

We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation where no musical score…

声音 · 计算机科学 2019-02-05 Sanna Wager , George Tzanetakis , Cheng-i Wang , Lijiang Guo , Aswin Sivaraman , Minje Kim

Automatic note-level transcription is considered one of the most challenging tasks in music information retrieval. The specific case of flamenco singing transcription poses a particular challenge due to its complex melodic progressions,…

声音 · 计算机科学 2016-11-17 Nadine Kroher , Emilia Gómez

This paper presents a novel supervised approach to detecting the chorus segments in popular music. Traditional approaches to this task are mostly unsupervised, with pipelines designed to target some quality that is assumed to define…

音频与语音处理 · 电气工程与系统科学 2021-04-22 Ju-Chiang Wang , Jordan B. L. Smith , Jitong Chen , Xuchen Song , Yuxuan Wang

Most of the previous approaches to lyrics-to-audio alignment used a pre-developed automatic speech recognition (ASR) system that innately suffered from several difficulties to adapt the speech model to individual singers. A significant…

声音 · 计算机科学 2020-10-29 Sungkyun Chang , Kyogu Lee

Music Information Retrieval (MIR) is a collaborative scientific study that help to build innovative information research themes, novel frameworks, and developing connected delivery mechanisms in addition to making the world's massive…

声音 · 计算机科学 2021-09-09 Shah Riya Chiragkumar

In this paper, we explore prosody transfer for audiobook generation under rather realistic condition where training DB is plain audio mostly from multiple ordinary people and reference audio given during inference is from professional and…

音频与语音处理 · 电气工程与系统科学 2020-05-22 Sunghee Jung , Hoirin Kim
‹ 上一页 1 2 3 10 下一页 ›