English
Related papers

Related papers: The Spotify Podcast Dataset

200 papers

Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such…

Sound · Computer Science 2024-03-08 Xuenan Xu , Xiaohang Xu , Zeyu Xie , Pingyue Zhang , Mengyue Wu , Kai Yu

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

Computation and Language · Computer Science 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions to advance this…

Sound · Computer Science 2026-03-18 Alejandro Paredes La Torre

Recent advances in text-to-speech (TTS) synthesis have significantly improved speech expressiveness and naturalness. However, most existing systems are tailored for single-speaker synthesis and fall short in generating coherent…

Radio remains a pervasive medium for mass information dissemination, with AM/FM stations reaching more Americans than either smartphone-based social networking or live television. Increasingly, radio broadcasts are also streamed online and…

Information Retrieval · Computer Science 2025-01-30 Govind Mittal , Sarthak Gupta , Shruti Wagle , Chirag Chopra , Anthony J DeMattee , Nasir Memon , Mustaque Ahamad , Chinmay Hegde

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

This work presents a user-centric recommendation framework, designed as a pipeline with four distinct, connected, and customizable phases. These phases are intended to improve explainability and boost user engagement. We have collected the…

Information Retrieval · Computer Science 2025-05-19 Jaime Ramirez Castillo , M. Julia Flores , Ann E. Nicholson

In this work, we showcase a cost-effective method for generating training data for speech processing tasks. First, we transcribe unlabeled speech using a state-of-the-art Automatic Speech Recognition (ASR) model. Next, we align generated…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-19 Taras Sereda

People rely on news to know what is happening around the world and inform their daily lives. In today's world, when the proliferation of fake news is rampant, having a large-scale and high-quality source of authentic news articles with the…

Computation and Language · Computer Science 2022-10-10 Rishabh Misra

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), containing 5,193 video…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Yidan Sun , Qin Chao , Yangfeng Ji , Boyang Li

It remains unknown whether personalized recommendations increase or decrease the diversity of content people consume. We present results from a randomized field experiment on Spotify testing the effect of personalized recommendations on…

Social and Information Networks · Computer Science 2020-03-19 David Holtz , Benjamin Carterette , Praveen Chandar , Zahra Nazari , Henriette Cramer , Sinan Aral

Affective computing aims to understand and model human emotions for computational systems. Within this field, speech emotion recognition (SER) focuses on predicting emotions conveyed through speech. While early SER systems relied on limited…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-25 Luz Martinez-Lucas , Pravin Mote , Abinay Reddy Naini , Mohammed Abdelwahab , Carlos Busso

Video podcast teasers are short videos that can be shared on social media platforms to capture interest in the full episodes of a video podcast. These teasers enable long-form podcasters to reach new audiences and gain new followers.…

Human-Computer Interaction · Computer Science 2024-05-10 Sitong Wang , Zheng Ning , Anh Truong , Mira Dontcheva , Dingzeyu Li , Lydia B. Chilton

Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization remains loosely defined. The…

Computation and Language · Computer Science 2025-10-20 Fabian Retkowski , Maike Züfle , Andreas Sudmann , Dinah Pfau , Shinji Watanabe , Jan Niehues , Alexander Waibel

The dynamic propagation of social media has broadened the reach of financial advisory content through podcast videos, yet extracting insights from lengthy, multimodal segments (30-40 minutes) remains challenging. We introduce FASTER…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Sarmistha Das , R E Zera Marveen Lyngkhoi , Sriparna Saha , Alka Maurya

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the…

Sound · Computer Science 2024-10-07 Gabriel Bibbó , Thomas Deacon , Arshdeep Singh , Mark D. Plumbley

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled…

Computation and Language · Computer Science 2024-06-04 Xinjian Li , Shinnosuke Takamichi , Takaaki Saeki , William Chen , Sayaka Shiota , Shinji Watanabe

Audio event detection is a widely studied audio processing task, with applications ranging from self-driving cars to healthcare. In-the-wild datasets such as Audioset have propelled research in this field. However, many efforts typically…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-16 Rajat Hebbar , Digbalay Bose , Krishna Somandepalli , Veena Vijai , Shrikanth Narayanan

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

Sound · Computer Science 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang
‹ Prev 1 4 5 6 7 8 10 Next ›