English
Related papers

Related papers: The Spotify Podcast Dataset

200 papers

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains.…

Computation and Language · Computer Science 2025-05-27 Dongqi Liu , Chenxi Whitehouse , Xi Yu , Louis Mahon , Rohit Saxena , Zheng Zhao , Yifu Qiu , Mirella Lapata , Vera Demberg

Our analysis reviews and visualizes the audio features and popularity of songs streamed on Spotify*. Our dataset, downloaded from Kaggle and originally sourced from Spotify API, consists of multiple Excel files containing information…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-08-06 Virginia Ochi , Ricardo Estrada , Teezal Gaji , Wendy Gadea , Emily Duong

More than two years after its outbreak, the COVID-19 pandemic continues to plague medical systems around the world, putting a strain on scarce resources, and claiming human lives. From the very beginning, various AI-based COVID-19 detection…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-23 Andreas Triantafyllopoulos , Anastasia Semertzidou , Meishu Song , Florian B. Pokorny , Björn W. Schuller

Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for…

Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses this issue with a meticulously annotated corpus of 12K…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Dongjie Yang , Suyuan Huang , Chengqiang Lu , Xiaodong Han , Haoxin Zhang , Yan Gao , Yao Hu , Hai Zhao

We present a unified multi-objective model for targeting both advertisements and promotions within the Spotify podcast ecosystem. Our approach addresses key challenges in personalization and cold-start initialization, particularly for new…

Information Retrieval · Computer Science 2026-01-06 Shivam Verma , Hannes Karlbom , Yu Zhao , Nick Topping , Vivian Chen , Kieran Stanley , Bharath Rengarajan

In the age of digital music streaming, playlists on platforms like Spotify have become an integral part of individuals' musical experiences. People create and publicly share their own playlists to express their musical tastes, promote the…

Cryptography and Security · Computer Science 2025-12-16 Pier Paolo Tricomi , Luca Pajola , Luca Pasa , Mauro Conti

The mood of a song is a highly relevant feature for exploration and recommendation in large collections of music. These collections tend to require automatic methods for predicting such moods. In this work, we show that listening-based…

Sound · Computer Science 2020-10-24 Filip Korzeniowski , Oriol Nieto , Matthew McCallum , Minz Won , Sergio Oramas , Erik Schmidt

The accelerating pace of scientific publishing makes it increasingly difficult for researchers to stay current. We present Paper Espresso, an open-source platform that automatically discovers, summarizes, and analyzes trending arXiv papers.…

Digital Libraries · Computer Science 2026-04-07 Mingzhe Du , Luu Anh Tuan , Dong Huang , See-kiong Ng

Film media is a rich form of artistic expression. Unlike photography, and short videos, movies contain a storyline that is deliberately complex and intricate in order to engage its audience. In this paper we present a large scale study…

Computer Vision and Pattern Recognition · Computer Science 2019-08-09 Paola Cascante-Bonilla , Kalpathy Sitaraman , Mengjia Luo , Vicente Ordonez

In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as information about the…

Computation and Language · Computer Science 2023-12-14 Ekaterina Garmash , Edgar Tanaka , Ann Clifton , Joana Correia , Sharmistha Jat , Winstead Zhu , Rosie Jones , Jussi Karlgren

With the explosive growth of livestream broadcasting, there is an urgent need for new summarization technology that enables us to create a preview of streamed content and tap into this wealth of knowledge. However, the problem is nontrivial…

Computation and Language · Computer Science 2021-09-14 Sangwoo Cho , Franck Dernoncourt , Tim Ganter , Trung Bui , Nedim Lipka , Walter Chang , Hailin Jin , Jonathan Brandt , Hassan Foroosh , Fei Liu

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

Every day we are surrounded by spoken dialog. This medium delivers rich diverse streams of information auditorily; however, systematically understanding dialog can often be non-trivial. Despite the pervasiveness of spoken dialog, automated…

Computation and Language · Computer Science 2021-08-24 Daniel Li , Thomas Chen , Albert Tung , Lydia Chilton

With the advancement in personal smart devices and pervasive network connectivity, users are no longer passive content consumers, but also contributors in producing new contents. This expansion in live services requires a detailed analysis…

Multimedia · Computer Science 2020-03-25 Emna Baccour , Aiman Erbad , Kashif Bilal , Amr Mohamed , Mohsen Guizani , Mounir Hamdi

The AI-Reporter represents a paradigmatic shift in scientific publication practice. This document demonstrates through a concrete case study how our system transforms academic presentations into publication-ready chapters -- in less than…

Digital Libraries · Computer Science 2025-08-01 Gerd Graßhoff

One important class of online videos is that of news broadcasts. Most news organisations provide near-immediate access to topical news broadcasts over the Internet, through RSS streams or podcasts. Until lately, technology has not made it…

Automated audio captioning is a cross-modal translation task that aims to generate natural language descriptions for given audio clips. This task has received increasing attention with the release of freely available datasets in recent…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Xinhao Mei , Xubo Liu , Mark D. Plumbley , Wenwu Wang

We release synth1B1, a multi-modal audio corpus consisting of 1 billion 4-second synthesized sounds, paired with the synthesis parameters used to generate them. The dataset is 100x larger than any audio dataset in the literature. We also…

Sound · Computer Science 2021-07-21 Joseph Turian , Jordie Shier , George Tzanetakis , Kirk McNally , Max Henry

Scientific news reports serve as a bridge, adeptly translating complex research articles into reports that resonate with the broader public. The automated generation of such narratives enhances the accessibility of scholarly insights. In…

Computation and Language · Computer Science 2024-12-11 Dongqi Liu , Yifan Wang , Jia Loy , Vera Demberg