English
Related papers

Related papers: Multi-Source Transformer Architectures for Audiovi…

200 papers

Visual-to-auditory sensory substitution devices can assist the blind in sensing the visual environment by translating the visual information into a sound pattern. To improve the translation quality, the task performances of the blind are…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Dong Wang , Xuelong Li , Feiping Nie , Qi Wang

A central problem in building effective sound event detection systems is the lack of high-quality, strongly annotated sound event datasets. For this reason, Task 4 of the DCASE 2024 challenge proposes learning from two heterogeneous…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-19 Florian Schmid , Paul Primus , Tobias Morocutti , Jonathan Greif , Gerhard Widmer

Previous DCASE challenges contributed to an increase in the performance of acoustic scene classification systems. State-of-the-art classifiers demand significant processing capabilities and memory which is challenging for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-10 Nagashree K. S. Rao , Nils Peters

Visible images offer rich texture details, while infrared images emphasize salient targets. Fusing these complementary modalities enhances scene understanding, particularly for advanced vision tasks under challenging conditions. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Beining Xu , Junxian Li

Deep convolution neural network has attracted many attentions in large-scale visual classification task, and achieves significant performance improvement compared to traditional visual analysis methods. In this paper, we explore many kinds…

Computer Vision and Pattern Recognition · Computer Science 2020-07-06 Feifei Huang , Jie Li , Xuelin Zhu

Scene recognition based on deep-learning has made significant progress, but there are still limitations in its performance due to challenges posed by inter-class similarities and intra-class dissimilarities. Furthermore, prior research has…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Amirhossein Aminimehr , Amirali Molaei , Erik Cambria

Audio scene classification, the problem of predicting class labels of audio scenes, has drawn lots of attention during the last several years. However, it remains challenging and falls short of accuracy and efficiency. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2018-05-21 Kele Xu , Dawei Feng , Haibo Mi , Boqing Zhu , Dezhi Wang , Lilun Zhang , Hengxing Cai , Shuwen Liu

Acoustic Scene Classification (ASC) is a challenging task, as a single scene may involve multiple events that contain complex sound patterns. For example, a cooking scene may contain several sound sources including silverware clinking,…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-20 Weimin Wang , Weiran Wang , Ming Sun , Chao Wang

Multi-label image classification is about predicting a set of class labels that can be considered as orderless sequential data. Transformers process the sequential data as a whole, therefore they are inherently good at set prediction. The…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Vacit Oguz Yazici , Joost van de Weijer , Longlong Yu

Image classification with small datasets has been an active research area in the recent past. However, as research in this scope is still in its infancy, two key ingredients are missing for ensuring reliable and truthful progress: a…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 L. Brigato , B. Barz , L. Iocchi , J. Denzler

This paper presents a task of audio-visual scene classification (SC) where input videos are classified into one of five real-life crowded scenes: 'Riot', 'Noise-Street', 'Firework-Event', 'Music-Event', and 'Sport-Atmosphere'. To this end,…

Computer Vision and Pattern Recognition · Computer Science 2021-12-20 Lam Pham , Dat Ngo , Phu X. Nguyen , Truong Hoang , Alexander Schindler

This technical report describes the IOA team's submission for TASK1A of DCASE2019 challenge. Our acoustic scene classification (ASC) system adopts a data augmentation scheme employing generative adversary networks. Two major classifiers, 1D…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-17 Hangting Chen , Zuozhen Liu , Zongming Liu , Pengyuan Zhang , Yonghong Yan

While text-to-video diffusion models have advanced significantly, creating coherent long-form content remains unreliable due to stochastic sampling artifacts. This necessitates generating multiple candidates, yet verifying them creates a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Daewon Yoon , Hyeongseok Lee , Wonsik Shin , Sangyu Han , Nojun Kwak

In recent years, vision transformers with text decoder have demonstrated remarkable performance on Scene Text Recognition (STR) due to their ability to capture long-range dependencies and contextual relationships with high learning…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Savas Ozkan , Andrea Maracani , Hyowon Kim , Sijun Cho , Eunchung Noh , Jeongwon Min , Jung Min Cho , Mete Ozay

Medical images commonly exhibit multiple abnormalities. Predicting them requires multi-class classifiers whose training and desired reliable performance can be affected by a combination of factors, such as, dataset size, data source,…

Image and Video Processing · Electrical Eng. & Systems 2021-11-16 Sivaramakrishnan Rajaraman , Ghada Zamzmi , Sameer Antani

In multiphase flow systems, classifying flow patterns is crucial to optimize fluid dynamics and enhance system efficiency. Current industrial methods and scientific laboratories mainly depend on techniques such as flow visualization using…

Machine Learning · Computer Science 2025-02-27 Nian Ran , Fayez M. Al-Alweet , Richard Allmendinger , Ahmad Almakhlafi

In this technical report, a low-complexity deep learning system for acoustic scene classification (ASC) is presented. The proposed system comprises two main phases: (Phase I) Training a teacher network; and (Phase II) training a student…

Sound · Computer Science 2023-05-17 Lam Pham , Dat Ngo , Cam Le , Anahid Jalali , Alexander Schindler

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation…

Artificial Intelligence · Computer Science 2025-01-16 Mathieu Lagrange , Junwon Lee , Modan Tailleur , Laurie M. Heller , Keunwoo Choi , Brian McFee , Keisuke Imoto , Yuki Okamoto

We introduce in this work an efficient approach for audio scene classification using deep recurrent neural networks. An audio scene is firstly transformed into a sequence of high-level label tree embedding feature vectors. The vector…

Sound · Computer Science 2017-06-06 Huy Phan , Philipp Koch , Fabrice Katzberg , Marco Maass , Radoslaw Mazur , Alfred Mertins

Pixel-level Scene Understanding is one of the fundamental problems in computer vision, which aims at recognizing object classes, masks and semantics of each pixel in the given image. Since the real-world is actually video-based rather than…

Image and Video Processing · Electrical Eng. & Systems 2023-06-06 Biao Wu , Shaoli Liu , Diankai Zhang , Chengjian Zheng , Si Gao , Xiaofeng Zhang , Ning Wang