中文
相关论文

相关论文: Beyond Equal-Length Snippets: How Long is Sufficie…

200 篇论文

The novelty of this study consists in a multi-modality approach to scene classification, where image and audio complement each other in a process of deep late fusion. The approach is demonstrated on a difficult classification problem,…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Jordan J. Bird , Diego R. Faria , Cristiano Premebida , Anikó Ekárt , George Vogiatzis

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Haytham M. Fayek , Anurag Kumar

Automatic speech recognition (ASR) models are normally trained to operate over single utterances, with a short duration of less than 30 seconds. This choice has been made in part due to computational constraints, but also reflects a common,…

音频与语音处理 · 电气工程与系统科学 2026-02-11 Robert Flynn , Anton Ragni

The identification of device brands and models plays a pivotal role in the realm of multimedia forensic applications. This paper presents a framework capable of identifying devices using audio, visual content, or a fusion of them. The…

机器学习 · 计算机科学 2024-06-27 Ioannis Tsingalis , Christos Korgialas , Constantine Kotropoulos

Image-based sequence recognition has been a long-standing research topic in computer vision. In this paper, we investigate the problem of scene text recognition, which is among the most important and challenging tasks in image-based…

计算机视觉与模式识别 · 计算机科学 2015-07-22 Baoguang Shi , Xiang Bai , Cong Yao

This paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multiple European cities, using the same equipment, in multiple…

音频与语音处理 · 电气工程与系统科学 2021-02-12 Shanshan Wang , Annamaria Mesaros , Toni Heittola , Tuomas Virtanen

We propose a new deep network for audio event recognition, called AENet. In contrast to speech, sounds coming from audio events may be produced by a wide variety of sources. Furthermore, distinguishing them often requires analyzing an…

多媒体 · 计算机科学 2017-01-05 Naoya Takahashi , Michael Gygli , Luc Van Gool

We introduce in this work an efficient approach for audio scene classification using deep recurrent neural networks. An audio scene is firstly transformed into a sequence of high-level label tree embedding feature vectors. The vector…

声音 · 计算机科学 2017-06-06 Huy Phan , Philipp Koch , Fabrice Katzberg , Marco Maass , Radoslaw Mazur , Alfred Mertins

Many applications of cross-modal music retrieval are related to connecting sheet music images to audio recordings. A typical and recent approach to this is to learn, via deep neural networks, a joint embedding space that correlates short…

声音 · 计算机科学 2023-09-22 Luis Carvalho , Gerhard Widmer

Audio scene classification, the problem of predicting class labels of audio scenes, has drawn lots of attention during the last several years. However, it remains challenging and falls short of accuracy and efficiency. Recently,…

计算机视觉与模式识别 · 计算机科学 2018-05-21 Kele Xu , Dawei Feng , Haibo Mi , Boqing Zhu , Dezhi Wang , Lilun Zhang , Hengxing Cai , Shuwen Liu

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

音频与语音处理 · 电气工程与系统科学 2022-09-27 Wim Boes , Hugo Van hamme

The use of multiple and semantically correlated sources can provide complementary information to each other that may not be evident when working with individual modalities on their own. In this context, multi-modal models can help producing…

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

计算机视觉与模式识别 · 计算机科学 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal…

计算机视觉与模式识别 · 计算机科学 2018-11-13 Bruno Korbar , Du Tran , Lorenzo Torresani

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

The sound of crashing waves, the roar of fast-moving cars -- sound conveys important information about the objects in our surroundings. In this work, we show that ambient sounds can be used as a supervisory signal for learning visual…

计算机视觉与模式识别 · 计算机科学 2017-12-21 Andrew Owens , Jiajun Wu , Josh H. McDermott , William T. Freeman , Antonio Torralba

Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale…

声音 · 计算机科学 2021-05-21 Sascha Hornauer , Ke Li , Stella X. Yu , Shabnam Ghaffarzadegan , Liu Ren

Recognizing acoustic events is an intricate problem for a machine and an emerging field of research. Deep neural networks achieve convincing results and are currently the state-of-the-art approach for many tasks. One advantage is their…

神经与进化计算 · 计算机科学 2016-03-21 Lars Hertel , Huy Phan , Alfred Mertins

Modern audio source separation techniques rely on optimizing sequence model architectures such as, 1D-CNNs, on mixture recordings to generalize well to unseen mixtures. Specifically, recent focus is on time-domain based architectures such…

‹ 上一页 1 2 3 10 下一页 ›