English
Related papers

Related papers: Rhapsody: A Dataset for Highlight Detection in Pod…

200 papers

Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses this issue with a meticulously annotated corpus of 12K…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Dongjie Yang , Suyuan Huang , Chengqiang Lu , Xiaodong Han , Haoxin Zhang , Yan Gao , Yao Hu , Hai Zhao

Audio captioning is the novel task of general audio content description using free text. It is an intermodal translation task (not speech-to-text), where a system accepts as an input an audio signal and outputs the textual description (i.e.…

Sound · Computer Science 2019-10-22 Konstantinos Drossos , Samuel Lipping , Tuomas Virtanen

Massive web datasets play a key role in the success of large vision-language models like CLIP and Flamingo. However, the raw web data is noisy, and existing filtering methods to reduce noise often come at the expense of data diversity. Our…

Machine Learning · Computer Science 2023-10-27 Thao Nguyen , Samir Yitzhak Gadre , Gabriel Ilharco , Sewoong Oh , Ludwig Schmidt

Highlight detection has the potential to significantly ease video browsing, but existing methods often suffer from expensive supervision requirements, where human viewers must manually identify highlights in training videos. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Bo Xiong , Yannis Kalantidis , Deepti Ghadiyaram , Kristen Grauman

While there is an abundance of popular writing targeted to podcast creators on how to speak in ways that engage their listeners, there has been little data-driven analysis of podcasts that relates linguistic style with listener engagement.…

Computation and Language · Computer Science 2021-06-15 Sravana Reddy , Marina Lazarova , Yongze Yu , Rosie Jones

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Matthew Gwilliam , Michael Cogswell , Meng Ye , Karan Sikka , Abhinav Shrivastava , Ajay Divakaran

Most pattern mining methods output a very large number of frequent patterns and isolating a small but relevant subset is a challenging problem of current interest in frequent pattern mining. In this paper we consider discovery of a small…

Databases · Computer Science 2014-10-14 A. Ibrahim , Shivakumar Sastry , P. S. Sastry

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

Audio-visual video highlight detection aims to automatically identify the most salient moments in videos by leveraging both visual and auditory cues. However, existing models often underutilize the audio modality, focusing on high-level…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-06 Seohyun Joo , Yoori Oh

Dialogue segmentation is a crucial task for dialogue systems allowing a better understanding of conversational texts. Despite recent progress in unsupervised dialogue segmentation methods, their performances are limited by the lack of…

Computation and Language · Computer Science 2023-10-17 Junfeng Jiang , Chengzhang Dong , Sadao Kurohashi , Akiko Aizawa

Whisper, despite being trained on 680K hours of web-scaled audio data, faces difficulty in recognising rare words like domain-specific terms, with a solution being contextual biasing through prompting. To improve upon this method, in this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-19 Yash Jogi , Vaibhav Aggarwal , Shabari S Nair , Yash Verma , Aayush Kubba

This paper introduces GigaSpeech, an evolving, multi-domain English speech recognition corpus with 10,000 hours of high quality labeled audio suitable for supervised training, and 40,000 hours of total audio suitable for semi-supervised and…

The goal of music highlight extraction is to get a short consecutive segment of a piece of music that provides an effective representation of the whole piece. In a previous work, we introduced an attention-based convolutional recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-27 Yu-Siang Huang , Szu-Yu Chou , Yi-Hsuan Yang

Establishing a good information retrieval system in popular mediums of entertainment is a quickly growing area of investigation for companies and researchers alike. We delve into the domain of information retrieval for podcasts. In…

Information Retrieval · Computer Science 2021-03-09 Abheesht Sharma , Harshit Pandey

Past studies in Sarcasm Detection mostly make use of Twitter datasets collected using hashtag-based supervision but such datasets are noisy in terms of labels and language. Furthermore, many tweets are replies to other tweets, and detecting…

Computation and Language · Computer Science 2022-12-13 Rishabh Misra

One of the biggest setbacks in traditional frequent pattern mining is that overwhelmingly many of the discovered patterns are redundant. A prototypical example of such redundancy is a freerider pattern where the pattern contains a true…

Data Structures and Algorithms · Computer Science 2019-02-05 Nikolaj Tatti

Audio tagging is the task of predicting the presence or absence of sound classes within an audio clip. Previous work in audio tagging focused on relatively small datasets limited to recognising a small number of sound classes. We…

Sound · Computer Science 2019-12-11 Qiuqiang Kong , Changsong Yu , Turab Iqbal , Yong Xu , Wenwu Wang , Mark D. Plumbley

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

Computation and Language · Computer Science 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

In recent years, there has been a notable increase in research on machine learning models for music retrieval and generation systems that are capable of taking natural language sentences as inputs. However, there is a scarcity of…

Computation and Language · Computer Science 2025-01-07 Takashi Harada , Takehiro Motomitsu , Katsuhiko Hayashi , Yusuke Sakai , Hidetaka Kamigaito

Audio captioning is an important research area that aims to generate meaningful descriptions for audio clips. Most of the existing research extracts acoustic features of audio clips as input to encoder-decoder and transformer architectures…

Sound · Computer Science 2022-04-20 Ayşegül Özkaya Eren , Mustafa Sert