中文
相关论文

相关论文: The Spotify Podcast Dataset

200 篇论文

We introduce SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original…

Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench, a benchmark…

计算与语言 · 计算机科学 2026-01-22 Chenning Xu , Mao Zheng , Mingyu Zheng , Mingyang Song

Evaluating personalized recommendations remains a central challenge, especially in long-form audio domains like podcasts, where traditional offline metrics suffer from exposure bias and online methods such as A/B testing are costly and…

Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music…

Scientific talks are a growing medium for disseminating research, and automatically identifying relevant literature that grounds or enriches a talk would be highly valuable for researchers and students alike. We introduce Reference…

计算与语言 · 计算机科学 2025-10-29 Frederik Broy , Maike Züfle , Jan Niehues

Sequence-to-sequence models have recently gained the state of the art performance in summarization. However, not too many large-scale high-quality datasets are available and almost all the available ones are mainly news articles with…

计算与语言 · 计算机科学 2018-10-23 Mahnaz Koupaee , William Yang Wang

MediaSum, a large-scale media interview dataset consisting of 463.6K transcripts with abstractive summaries. To create this dataset, we collect interview transcripts from NPR and CNN and employ the overview and topic descriptions as…

计算与语言 · 计算机科学 2021-03-15 Chenguang Zhu , Yang Liu , Jie Mei , Michael Zeng

Narrative summarization aims to produce a distilled version of a narrative to describe its most salient events and characters. Summarizing a narrative is challenging as it requires an understanding of event causality and character…

计算与语言 · 计算机科学 2023-06-29 Chao Zhao , Faeze Brahman , Kaiqiang Song , Wenlin Yao , Dian Yu , Snigdha Chaturvedi

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximately 36.7k hours of English speech for 22.05 kHz training,…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Ryan Langman , Xuesong Yang , Paarth Neekhara , Shehzeen Hussain , Edresson Casanova , Evelina Bakhturina , Jason Li

We study how participation in collective action is articulated in podcast discussions, using the Black Lives Matter (BLM) movement as a case study. While research on collective action discourse has primarily focused on text-based content,…

社会与信息网络 · 计算机科学 2025-09-17 Theodora Moldovan , Arianna Pera , Davide Vega , Luca Maria Aiello

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

计算机视觉与模式识别 · 计算机科学 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

Compared to search engine result pages (SERPs), AI-generated podcasts represent a relatively new and relatively more passive modality of information consumption, delivering narratives in a naturally engaging format. As these two media…

信息检索 · 计算机科学 2026-01-19 Junjie Wang , Gaole He , Alisa Rieger , Ujwal Gadiraju

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the…

声音 · 计算机科学 2023-12-27 Haoxu Wang , Fan Yu , Xian Shi , Yuezhang Wang , Shiliang Zhang , Ming Li

Between January 2017 and January 2021, thousands of local news sources in the United States reported on over 42,000 protests about topics such as civil rights, immigration, guns, and the environment. Given the vast number of local…

计算与语言 · 计算机科学 2021-02-02 Tommy Leung , L. Nathan Perkins

A large and growing amount of speech content in real-life scenarios is being recorded on consumer-grade devices in uncontrolled environments, resulting in degraded speech quality. Transforming such low-quality device-degraded speech into…

音频与语音处理 · 电气工程与系统科学 2022-03-23 Haoyu Li , Junichi Yamagishi

Paralinguistic sounds, like laughter and sighs, are crucial for synthesizing more realistic and engaging speech. However, existing methods typically depend on proprietary datasets, while publicly available resources often suffer from…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Bingsong Bai , Qihang Lu , Wenbing Yang , Zihan Sun , Yueran Hou , Peilei Jia , Songbai Pu , Ruibo Fu , Yingming Gao , Ya Li , Jun Gao

For over a decade, TV series have been drawing increasing interest, both from the audience and from various academic fields. But while most viewers are hooked on the continuous plots of TV serials, the few annotated datasets available to…

信息检索 · 计算机科学 2021-01-19 Xavier Bost , Vincent Labatut , Georges Linares

Podcasts are conversational in nature and speaker changes are frequent -- requiring speaker diarization for content understanding. We propose an unsupervised technique for speaker diarization without relying on language-specific components.…

计算与语言 · 计算机科学 2022-07-27 M. Iftekhar Tanveer , Diego Casabuena , Jussi Karlgren , Rosie Jones

The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released…

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited,…