English
Related papers

Related papers: Clotho: An Audio Captioning Dataset

200 papers

Disentangling conversations mixed together in a single stream of messages is a difficult task, made harder by the lack of large manually annotated datasets. We created a new dataset of 77,563 messages manually annotated with reply-structure…

Automatic image caption generation aims to produce an accurate description of an image in natural language automatically. However, Bangla, the fifth most widely spoken language in the world, is lagging considerably in the research and…

Computation and Language · Computer Science 2018-09-10 Motiur Rahman , Nabeel Mohammed , Nafees Mansoor , Sifat Momen

Large Language Models (LLMs) have shown immense potential in multimodal applications, yet the convergence of textual and musical domains remains not well-explored. To address this gap, we present MusiLingo, a novel system for music caption…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-03 Zihao Deng , Yinghao Ma , Yudong Liu , Rongchen Guo , Ge Zhang , Wenhu Chen , Wenhao Huang , Emmanouil Benetos

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

Audio classification is the task of identifying the sound categories that are associated with a given audio signal. This paper presents an investigation on large-scale audio classification based on the recently released AudioSet database.…

Sound · Computer Science 2018-10-31 Yuzhong Wu , Tan Lee

Lifelogging cameras capture everyday life from a first-person perspective, but generate so much data that it is hard for users to browse and organize their image collections effectively. In this paper, we propose to use automatic image…

Computer Vision and Pattern Recognition · Computer Science 2016-08-15 Chenyou Fan , David J. Crandall

While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when…

Sound · Computer Science 2025-01-07 Xiquan Li , Wenxi Chen , Ziyang Ma , Xuenan Xu , Yuzhe Liang , Zhisheng Zheng , Qiuqiang Kong , Xie Chen

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tasks, hindering the…

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

We present MeetDot, a videoconferencing system with live translation captions overlaid on screen. The system aims to facilitate conversation between people who speak different languages, thereby reducing communication barriers between…

Computation and Language · Computer Science 2021-09-21 Arkady Arkhangorodsky , Christopher Chu , Scot Fang , Yiqi Huang , Denglin Jiang , Ajay Nagesh , Boliang Zhang , Kevin Knight

Matching raw audio signals with textual descriptions requires understanding the audio's content and the description's semantics and then drawing connections between the two modalities. This paper investigates a hybrid retrieval system that…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-03 Paul Primus , Gerhard Widmer

Music captioning has gained significant attention in the wake of the rising prominence of streaming media platforms. Traditional approaches often prioritize either the audio or lyrics aspect of the music, inadvertently ignoring the…

Sound · Computer Science 2023-10-24 Zihao He , Weituo Hao , Wei-Tsung Lu , Changyou Chen , Kristina Lerman , Xuchen Song

The diverse nature, scale, and specificity of podcasts present a unique challenge to content discovery systems. Listeners often rely on text descriptions of episodes provided by the podcast creators to discover new content. Some factors…

Computation and Language · Computer Science 2020-09-23 Aneesh Vartakavi , Amanmeet Garg

Recent advances in language and speech modelling have made it possible to build autonomous voice assistants that understand and generate human dialogue in real time. These systems are increasingly being deployed in domains such as customer…

Artificial Intelligence · Computer Science 2025-09-08 Krittanon Kaewtawee , Wachiravit Modecrua , Krittin Pachtrachai , Touchapon Kraisingkorn

In this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-05 Hokuto Munakata , Taichi Nishimura , Shota Nakada , Tatsuya Komatsu

This paper introduces the Voices Obscured In Complex Environmental Settings (VOICES) corpus, a freely available dataset under Creative Commons BY 4.0. This dataset will promote speech and signal processing research of speech recorded by…

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

We proposed Audio Difference Captioning (ADC) as a new extension task of audio captioning for describing the semantic differences between input pairs of similar but slightly different audio clips. The ADC solves the problem that…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-24 Daiki Takeuchi , Yasunori Ohishi , Daisuke Niizumi , Noboru Harada , Kunio Kashino

Controllable image captioning models generate human-like image descriptions, enabling some kind of control over the generated captions. This paper focuses on controlling the caption length, i.e. a short and concise description or a long and…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Elad Hirsch , Ayellet Tal

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal