中文
相关论文

相关论文: The Spotify Podcast Dataset

200 篇论文

The world today is experiencing an abundance of music like no other time, and attempts to group music into clusters have become increasingly prevalent. Common standards for grouping music were songs, artists, and genres, with artists or…

人机交互 · 计算机科学 2021-03-01 Seokgi Kim , Jihye Park , Kihong Seong , Namwoo Cho , Junho Min , Hwajung Hong

Videos and Podcasts have established themselves as the medium of choice for civic dissemination, but also as carriers of misinformation. The emerging Science Communication Knowledge Infrastructure (SciCom KI), which curates these…

数字图书馆 · 计算机科学 2026-02-05 Tim Wittenborg , Niklas Stehr , Oliver Karras , Sören Auer

Prosody -- the melody of speech -- conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expressed by prosody alone…

计算与语言 · 计算机科学 2025-12-19 Aditya Yadavalli , Tiago Pimentel , Tamar I Regev , Ethan Wilcox , Alex Warstadt

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing…

The era of live-broadcast is back but with two major changes. First, unlike traditional TV broadcasts, content is now streamed over the Internet enabling it to reach a wider audience. Second, due to various user-generated content platforms…

社会与信息网络 · 计算机科学 2018-03-08 Aravindh Raman , Gareth Tyson , Nishanth Sastry

This paper introduces the HumTrans dataset, which is publicly available and primarily designed for humming melody transcription. The dataset can also serve as a foundation for downstream tasks such as humming melody based music generation.…

声音 · 计算机科学 2023-10-18 Shansong Liu , Xu Li , Dian Li , Ying Shan

Large Language Models are increasingly being deployed to extract structured data from unstructured and semi-structured sources: parsing invoices, medical records, and converting PDF documents to database entries. Yet existing benchmarks for…

计算与语言 · 计算机科学 2026-04-29 Abhinav Kumar Singh , Harsha Vardhan Khurdula , Yoeven D Khemlani , Vineet Agarwal

Phonetics is the scientific field concerned with the study of how speech is produced, heard and perceived. It abounds with data, such as acoustic speech recordings, neuroimaging data, or articulatory data. In this paper, we provide an…

Existing Existing automatic audio generation methods struggle to generate podcast-like audio programs effectively. The key challenges lie in in-depth content generation, appropriate and expressive voice production. This paper proposed…

声音 · 计算机科学 2025-03-04 Yujia Xiao , Lei He , Haohan Guo , Fenglong Xie , Tan Lee

We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttural, nasal, pained) have been explored in small-scale…

音频与语音处理 · 电气工程与系统科学 2025-09-26 Anuj Diwan , Zhisheng Zheng , David Harwath , Eunsol Choi

The production of media content has undergone tremendous changes in recent years. Multiple daily content updates are just as common for some platforms as is processing the provided content specifically for their target audiences. Such…

多媒体 · 计算机科学 2024-07-30 Alexander Weller , Werner Bleisteiner , Christian Hufnagel , Michael Iber

We introduce the Merkel Podcast Corpus, an audio-visual-text corpus in German collected from 16 years of (almost) weekly Internet podcasts of former German chancellor Angela Merkel. To the best of our knowledge, this is the first single…

计算与语言 · 计算机科学 2022-05-25 Debjoy Saha , Shravan Nayak , Timo Baumann

We present INDICVOICES, a dataset of natural and spontaneous speech containing a total of 7348 hours of read (9%), extempore (74%) and conversational (17%) audio from 16237 speakers covering 145 Indian districts and 22 languages. Of these…

Conversational music recommendation (CMR) research currently faces a tradeoff between authentic dialogue corpora that are limited in scale and synthesized corpora that scale up but whose conversations are artificially constructed rather…

信息检索 · 计算机科学 2026-05-12 Haven Kim , Julian McAuley

Commonly music has an obvious hierarchical structure, especially for the singing parts which usually act as the main melody in pop songs. However, most of the current singing annotation datasets only record symbolic information of music…

声音 · 计算机科学 2022-10-03 Xiao Fu , Xin Yuan , Jinglu Hu

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the…

音频与语音处理 · 电气工程与系统科学 2026-03-26 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Junbo Zhang , Jian Luan

Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack generalization to the full spectrum of human diversity in ethnicity, language, and age…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Shunian Chen , Hejin Huang , Yexin Liu , Zihan Ye , Pengcheng Chen , Chenghao Zhu , Michael Guan , Rongsheng Wang , Junying Chen , Guanbin Li , Ser-Nam Lim , Harry Yang , Benyou Wang

Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio, different people may perceive the same audio differently, resulting in caption disparities (i.e., one audio may correlate to…

声音 · 计算机科学 2022-04-19 Yiming Zhang , Hong Yu , Ruoyi Du , Zhanyu Ma , Yuan Dong

Speech is a hierarchical collection of text, prosody, emotions, dysfluencies, etc. Automatic transcription of speech that goes beyond text (words) is an underexplored problem. We focus on transcribing speech along with non-fluencies…

音频与语音处理 · 电气工程与系统科学 2024-12-03 Jiachen Lian , Xuanru Zhou , Zoe Ezzes , Jet Vonk , Brittany Morin , David Baquirin , Zachary Mille , Maria Luisa Gorno Tempini , Gopala Krishna Anumanchipalli