English
Related papers

Related papers: Mapping the Podcast Ecosystem with the Structured …

200 papers

OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podcasts, talk shows, teleconferences, and other conversations.…

Computation and Language · Computer Science 2025-09-08 Wei Chu , Yuanzhe Dong , Ke Tan , Dong Han , Xavier Menendez-Pidal , Ruchao Fan , Chenfeng Miao , Chanwoo Kim , Bhiksha Raj , Rita Singh

The consumption of podcast media has been increasing rapidly. Due to the lengthy nature of podcast episodes, users often carefully select which ones to listen to. Although episode descriptions aid users by providing a summary of the entire…

Information Retrieval · Computer Science 2023-07-26 Andrew Aquilina , Sean Diacono , Panagiotis Papapetrou , Maria Movin

We present SwissGPC v1.0, the first mid-to-large-scale corpus of spontaneous Swiss German speech, developed to support research in ASR, TTS, dialect identification, and related fields. The dataset consists of links to talk shows and…

Computation and Language · Computer Science 2025-09-25 Samuel Stucki , Mark Cieliebak , Jan Deriu

We introduce RadioTalk, a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019. The corpus is intended for use by researchers in the fields of natural…

Computation and Language · Computer Science 2019-09-18 Doug Beeferman , William Brannon , Deb Roy

Podcasts have recently shown a rapid rise in popularity. Summarization of podcast transcripts is of practical benefit to both content providers and consumers. It helps consumers to quickly decide whether they will listen to the podcasts and…

Computation and Language · Computer Science 2022-03-23 Kaiqiang Song , Chen Li , Xiaoyang Wang , Dong Yu , Fei Liu

Everyday speech conveys far more than words, it reflects who we are, how we feel, and the circumstances surrounding our interactions. Yet, most existing speech datasets are acted, limited in scale, and fail to capture the expressive…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-04 Zongyang Du , Shreeram Suresh Chandra , Ismail Rasim Ulgen , Aurosweta Mahapatra , Ali N. Salman , Carlos Busso , Berrak Sisman

The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and…

Listening to audio content, such as podcasts and audiobooks, is one way for people to engage with knowledge. Listening affords people more mobility than reading by seeing, thereby broadening their learning opportunities. This study explores…

Human-Computer Interaction · Computer Science 2025-01-23 Yuchi Yahagi , Rintaro Chujo , Yuga Harada , Changyo Han , Kohei Sugiyama , Takeshi Naemura

Podcast episodes often contain material extraneous to the main content, such as advertisements, interleaved within the audio and the written descriptions. We present classifiers that leverage both textual and listening patterns in order to…

Computation and Language · Computer Science 2021-06-15 Sravana Reddy , Yongze Yu , Aasish Pappu , Aswin Sivaraman , Rezvaneh Rezapour , Rosie Jones

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

Computation and Language · Computer Science 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

News podcasts are a popular medium to stay informed and dive deep into news topics. Today, most podcasts are handcrafted by professionals. In this work, we advance the state-of-the-art in automatically generated podcasts, making use of…

Human-Computer Interaction · Computer Science 2022-02-16 Philippe Laban , Elicia Ye , Srujay Korlakunta , John Canny , Marti A. Hearst

The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models.…

Computation and Language · Computer Science 2024-06-24 Georgios Paraskevopoulos , Chara Tsoukala , Athanasios Katsamanis , Vassilis Katsouros

We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and evaluating (deep learning) source separation models. To that…

Sound · Computer Science 2022-07-18 Nicolás Schmidt , Jordi Pons , Marius Miron

Voice conversion (VC) research traditionally depends on scripted or acted speech, which lacks the natural spontaneity of real-life conversations. While natural speech data is limited for VC, our study focuses on filling in this gap. We…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 Ali N. Salman , Zongyang Du , Shreeram Suresh Chandra , Ismail Rasim Ulgen , Carlos Busso , Berrak Sisman

Podcast script generation requires LLMs to synthesize structured, context-grounded dialogue from diverse inputs, yet systematic evaluation resources for this task remain limited. To bridge this gap, we introduce PodBench, a benchmark…

Computation and Language · Computer Science 2026-01-22 Chenning Xu , Mao Zheng , Mingyu Zheng , Mingyang Song

Discovering and evaluating long-form talk content such as videos and podcasts poses a significant challenge for users, as it requires a considerable time investment. Previews offer a practical solution by providing concise snippets that…

Information Retrieval · Computer Science 2025-06-04 Winstead Zhu , Ann Clifton , Azin Ghazimatin , Edgar Tanaka , Edward Ronan

Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited,…

This paper contains the description of our submissions to the summarization task of the Podcast Track in TREC (the Text REtrieval Conference) 2020. The goal of this challenge was to generate short, informative summaries that contain the key…

Computation and Language · Computer Science 2021-04-09 Rezvaneh Rezapour , Sravana Reddy , Ann Clifton , Rosie Jones

Distinguishing scripted from spontaneous speech is an essential tool for better understanding how speech styles influence speech processing research. It can also improve recommendation systems and discovery experiences for media users…

Computation and Language · Computer Science 2024-12-17 Shahar Elisha , Andrew McDowell , Mariano Beguerisse-Díaz , Emmanouil Benetos

Audio captioning is a novel field of multi-modal translation and it is the task of creating a textual description of the content of an audio signal (e.g. "people talking in a big room"). The creation of a dataset for this task requires a…

Sound · Computer Science 2019-07-23 Samuel Lipping , Konstantinos Drossos , Tuomas Virtanen