English
Related papers

Related papers: Rhapsody: A Dataset for Highlight Detection in Pod…

200 papers

The goal of video moment retrieval and highlight detection is to identify specific segments and highlights based on a given text query. With the rapid growth of video content and the overlap between these tasks, recent works have addressed…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Sung Jin Um , Dongjin Kim , Sangmin Lee , Jung Uk Kim

Audio-language models (ALMs) generate linguistic descriptions of sound-producing events and scenes. Advances in dataset creation and computational power have led to significant progress in this domain. This paper surveys 69 datasets used to…

Sound · Computer Science 2025-02-10 Gijs Wijngaard , Elia Formisano , Michele Esposito , Michel Dumontier

Highlight detection models are typically trained to identify cues that make visual content appealing or interesting for the general public, with the objective of reducing a video to such moments. However, the "interestingness" of a video…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 Ana García del Molino , Michael Gygli

We propose Lighthouse, a user-friendly library for reproducible video moment retrieval and highlight detection (MR-HD). Although researchers proposed various MR-HD approaches, the research community holds two main issues. The first is a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Taichi Nishimura , Shota Nakada , Hokuto Munakata , Tatsuya Komatsu

The task of audio captioning is similar in essence to tasks such as image and video captioning. However, it has received much less attention. We propose three desiderata for captioning audio -- (i) fluency of the generated text, (ii)…

Sound · Computer Science 2023-09-08 Tal Shaharabany , Ariel Shaulov , Lior Wolf

Imagine sitting in a presentation, trying to follow the speaker while simultaneously scanning the slides for relevant information. While the entire slide is visible, identifying the relevant regions can be challenging. As you focus on one…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Megha Mariam K M , C. V. Jawahar

In this paper, we test the hypothesis that interesting events in unstructured videos are inherently audiovisual. We combine deep image representations for object recognition and scene understanding with representations from an audiovisual…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Karel Mundnich , Alexandra Fenster , Aparna Khare , Shiva Sundaram

The time and expense required to collect and label audio data has been a prohibitive factor in the availability of domain specific audio datasets. As the predictive specificity of a classifier depends on the specificity of the labels it is…

Sound · Computer Science 2023-11-14 Blake Downward , Jon Nordby

Modern noise-cancelling headphones have significantly improved users' auditory experiences by removing unwanted background noise, but they can also block out sounds that matter to users. Machine learning (ML) models for sound event…

Machine Learning · Computer Science 2024-06-18 N Shashaank , Berker Banar , Mohammad Rasool Izadi , Jeremy Kemmerer , Shuo Zhang , Chuan-Che Huang

In music and speech, meaning is derived at multiple levels of context. Affect, for example, can be inferred both by a short sound token and by sonic patterns over a longer temporal window such as an entire recording. In this letter, we…

Sound · Computer Science 2022-09-12 Camille Noufi , Prateek Verma

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a…

Sound · Computer Science 2022-02-15 Ke Chen , Xingjian Du , Bilei Zhu , Zejun Ma , Taylor Berg-Kirkpatrick , Shlomo Dubnov

The short-form videos have explosive popularity and have dominated the new social media trends. Prevailing short-video platforms,~\textit{e.g.}, Kuaishou (Kwai), TikTok, Instagram Reels, and YouTube Shorts, have changed the way we consume…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Wentao Zhu , Yufang Huang , Xiufeng Xie , Wenxian Liu , Jincan Deng , Debing Zhang , Zhangyang Wang , Ji Liu

Early detection of depression from online social media posts holds promise for providing timely mental health interventions. In this work, we present a high-quality, expert-annotated dataset of 1,017 social media posts labeled with…

Computation and Language · Computer Science 2025-07-29 Prajval Bolegave , Pushpak Bhattacharya

Everyday sound recognition aims to infer types of sound events in audio streams. While many works succeeded in training models with high performance in a fully-supervised manner, they are still restricted to the demand of large quantities…

Sound · Computer Science 2022-12-20 Jinhua Liang , Huy Phan , Emmanouil Benetos

This paper aims for event recognition when video examples are scarce or even completely absent. The key in such a challenging setting is a semantic video representation. Rather than building the representation from individual attribute…

Computer Vision and Pattern Recognition · Computer Science 2015-11-10 Amirhossein Habibian , Thomas Mensink , Cees G. M. Snoek

We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2 million captions, collected through crowdsourcing rating data for The New Yorker's weekly cartoon caption…

Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to this as audio…

Multimedia · Computer Science 2024-01-19 Taichi Nishimura , Shota Nakada , Masayoshi Kondo

Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descriptions, and practical value in various applications, it has…

Sound · Computer Science 2025-09-09 Pedro Ramoneda , Pablo Alonso-Jiménez , Sergio Oramas , Xavier Serra , Dmitry Bogdanov

The development of video game streaming has grown rapidly, with major platforms such as YouTube and Twitch using different codecs. To support quality assessment models that work consistently across any codec, it is necessary to have access…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Rajesh Sureddi , Shreshth Saini , Avinab Saha , Alan C. Bovik

In this work, we study the association between song lyrics and mood through a data-driven analysis. Our data set consists of nearly one million songs, with song-mood associations derived from user playlists on the Spotify streaming…

Multimedia · Computer Science 2022-07-13 Shahrzad Naseri , Sravana Reddy , Joana Correia , Jussi Karlgren , Rosie Jones
‹ Prev 1 4 5 6 7 8 10 Next ›