Recognizing Film Entities in Podcasts
Information Retrieval
2018-09-25 v1 Computation and Language
Abstract
In this paper, we propose a Named Entity Recognition (NER) system to identify film titles in podcast audio. Taking inspiration from NER systems for noisy text in social media, we implement a two-stage approach that is robust to computer transcription errors and does not require significant computational expense to accommodate new film titles/releases. Evaluating on a diverse set of podcasts, we demonstrate more than a 20% increase in F1 score across three baseline approaches when combining fuzzy-matching with a linear model aware of film-specific metadata.
Cite
@article{arxiv.1809.08711,
title = {Recognizing Film Entities in Podcasts},
author = {Ahmet Salih Gundogdu and Arjun Sanghvi and Keith Harrigian},
journal= {arXiv preprint arXiv:1809.08711},
year = {2018}
}
Comments
4 pages, 1 figure. To appear in Proceedings of 2018 KDD Workshop on Machine Learning and Data Mining for Podcasts, August 2018, London, UK