中文
相关论文

相关论文: Clotho: An Audio Captioning Dataset

200 篇论文

Computational sign language research lacks the large-scale datasets that enables the creation of useful reallife applications. To date, most research has been limited to prototype systems on small domains of discourse, e.g. weather…

计算机视觉与模式识别 · 计算机科学 2021-05-07 Necati Cihan Camgoz , Ben Saunders , Guillaume Rochette , Marco Giovanelli , Giacomo Inches , Robin Nachtrab-Ribback , Richard Bowden

Captioning has attracted much attention in image and video understanding while a small amount of work examines audio captioning. This paper contributes a Mandarin-annotated dataset for audio captioning within a car scene. A sentence-level…

声音 · 计算机科学 2020-10-26 Xuenan Xu , Heinrich Dinkel , Mengyue Wu , Kai Yu

Audio editing aims to manipulate audio content based on textual descriptions, supporting tasks such as adding, removing, or replacing audio events. Despite recent progress, the lack of high-quality benchmark datasets and comprehensive…

声音 · 计算机科学 2026-02-03 Yuhang Jia , Hui Wang , Xin Nie , Yujie Guo , Lianru Gao , Yong Qin

Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments paired with text. In an effort to minimize the annotation…

计算机视觉与模式识别 · 计算机科学 2023-07-13 Yongrae Jo , Seongyun Lee , Aiden SJ Lee , Hyunji Lee , Hanseok Oh , Minjoon Seo

CAPTCHAs are employed as a security measure to differentiate human users from bots. A new sound-based CAPTCHA is proposed in this paper, which exploits the gaps between human voice and synthetic voice rather than relays on the auditory…

密码学与安全 · 计算机科学 2013-06-13 Haichang Gao , Honggang Liu , Dan Yao , Xiyang Liu , Uwe Aickelin

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label…

声音 · 计算机科学 2025-07-17 Jaesung Huh , Jacob Chalk , Evangelos Kazakos , Dima Damen , Andrew Zisserman

In recent years, there has been a notable increase in research on machine learning models for music retrieval and generation systems that are capable of taking natural language sentences as inputs. However, there is a scarcity of…

计算与语言 · 计算机科学 2025-01-07 Takashi Harada , Takehiro Motomitsu , Katsuhiko Hayashi , Yusuke Sakai , Hidetaka Kamigaito

Standard image captioning tasks such as COCO and Flickr30k are factual, neutral in tone and (to a human) state the obvious (e.g., "a man playing a guitar"). While such tasks are useful to verify that a machine understands the content of an…

计算机视觉与模式识别 · 计算机科学 2019-03-21 Kurt Shuster , Samuel Humeau , Hexiang Hu , Antoine Bordes , Jason Weston

Conversational memory is the process by which humans encode, retain and retrieve verbal, non-verbal and contextual information from a conversation. Since human memory is selective, differing recollections of the same events can lead to…

计算与语言 · 计算机科学 2024-10-16 Maria Tsfasman , Bernd Dudzik , Kristian Fenech , Andras Lorincz , Catholijn M. Jonker , Catharine Oertel

Mainstream Audio Analytics models are trained to learn under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled…

声音 · 计算机科学 2022-06-13 Benjamin Elizalde , Soham Deshmukh , Mahmoud Al Ismail , Huaming Wang

This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, which has been used for training and evaluating automatic…

声音 · 计算机科学 2019-04-08 Heiga Zen , Viet Dang , Rob Clark , Yu Zhang , Ron J. Weiss , Ye Jia , Zhifeng Chen , Yonghui Wu

Emotion is essential in spoken communication, yet most existing frameworks in speech emotion modeling rely on predefined categories or low-dimensional continuous attributes, which offer limited expressive capacity. Recent advances in speech…

音频与语音处理 · 电气工程与系统科学 2026-04-07 Tianhua Qi , Wenming Zheng , Björn W. Schuller , Zhaojie Luo , Haizhou Li

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

计算机视觉与模式识别 · 计算机科学 2022-01-25 Arjun R. Akula , Song-Chun Zhu

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remove these background sounds using speech enhancement or train…

音频与语音处理 · 电气工程与系统科学 2022-02-04 Chaitanya Narisetty , Emiru Tsunoo , Xuankai Chang , Yosuke Kashiwagi , Michael Hentschel , Shinji Watanabe

Describes an audio dataset of spoken words designed to help train and evaluate keyword spotting systems. Discusses why this task is an interesting challenge, and why it requires a specialized dataset that is different from conventional…

计算与语言 · 计算机科学 2018-04-11 Pete Warden

We introduce the Situated Corpus Of Understanding Transactions (SCOUT), a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration. The corpus was constructed from multiple Wizard-of-Oz experiments…

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text --…

计算机视觉与模式识别 · 计算机科学 2021-11-23 Karan Desai , Gaurav Kaul , Zubin Aysola , Justin Johnson

At present, Text-to-speech (TTS) systems that are trained with high-quality transcribed speech data using end-to-end neural models can generate speech that is intelligible, natural, and closely resembles human speech. These models are…

计算与语言 · 计算机科学 2023-03-02 Ajinkya Kulkarni , Atharva Kulkarni , Sara Abedalmonem Mohammad Shatnawi , Hanan Aldarmaki

We introduce Diffusion-based Audio Captioning (DAC), a non-autoregressive diffusion model tailored for diverse and efficient audio captioning. Although existing captioning models relying on language backbones have achieved remarkable…

计算与语言 · 计算机科学 2025-06-03 Manjie Xu , Chenxing Li , Xinyi Tu , Yong Ren , Ruibo Fu , Wei Liang , Dong Yu

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a…