English
Related papers

Related papers: A Toolchain for Comprehensive Audio/Video Analysis…

200 papers

The thud of a bouncing ball, the onset of speech as lips open -- when visual and audio events occur together, it suggests that there might be a common, underlying event that produced both signals. In this paper, we argue that the visual and…

Computer Vision and Pattern Recognition · Computer Science 2018-10-10 Andrew Owens , Alexei A. Efros

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching…

This paper addresses the challenge of automated violence detection in video frames captured by surveillance cameras, specifically focusing on classifying scenes as "fight" or "non-fight." This task is critical for enhancing unmanned…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Gourang Pathak , Abhay Kumar , Sannidhya Rawat , Shikha Gupta

Deep learning has recently achieved very promising results in a wide range of areas such as computer vision, speech recognition and natural language processing. It aims to learn hierarchical representations of data by using deep…

Computer Vision and Pattern Recognition · Computer Science 2015-12-11 Li Wang , Dennis Sng

We present an audio-visual multimodal approach for the task of zeroshot learning (ZSL) for classification and retrieval of videos. ZSL has been studied extensively in the recent past but has primarily been limited to visual modality and to…

Computer Vision and Pattern Recognition · Computer Science 2019-10-22 Kranti Kumar Parida , Neeraj Matiyali , Tanaya Guha , Gaurav Sharma

Automatic Deception Detection has been a hot research topic for a long time, using machine learning and deep learning to automatically detect deception, brings new light to this old field. In this paper, we proposed a voting-based method…

Machine Learning · Computer Science 2024-03-18 Lana Touma , Mohammad Al Horani , Manar Tailouni , Anas Dahabiah , Khloud Al Jallad

The seen birds twitter, the running cars accompany with noise, etc. These naturally audiovisual correspondences provide the possibilities to explore and understand the outside world. However, the mixed multiple objects and sounds make it…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Feiping Nie , Xuelong Li

With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with…

Sound · Computer Science 2024-09-17 Anna Wang , Da Liu , Zhiyu Zhang , Shengqiang Liu , Jie Gao , Yali Li

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mohamed Eltahir , Osamah Sarraj , Mohammed Bremoo , Mohammed Khurd , Abdulrahman Alfrihidi , Taha Alshatiri , Mohammad Almatrafi , Tanveer Hussain

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

Multimedia · Computer Science 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Despite the abundance of current researches working on the sentiment analysis from videos and audios, finding the best model that gives the highest accuracy rate is still considered a challenge for researchers in this field. The main…

Sound · Computer Science 2024-12-13 Antonio Fernandez , Suzan Awinat

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Chao Huang , Ruohan Gao , J. M. F. Tsang , Jan Kurcius , Cagdas Bilen , Chenliang Xu , Anurag Kumar , Sanjeel Parekh

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

Computers and Society · Computer Science 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

auDeep is a Python toolkit for deep unsupervised representation learning from acoustic data. It is based on a recurrent sequence to sequence autoencoder approach which can learn representations of time series data by taking into account…

Deep learning, with its robust aotomatic feature extraction capabilities, has demonstrated significant success in audio signal processing. Typically, these methods rely on static, pre-collected large-scale datasets for training, performing…

Sound · Computer Science 2024-12-19 Qisheng Xu , Yulin Sun , Yi Su , Qian Zhu , Xiaoyi Tan , Hongyu Wen , Zijian Gao , Kele Xu , Yong Dou , Dawei Feng

Human brain is continuously inundated with the multisensory information and their complex interactions coming from the outside world at any given moment. Such information is automatically analyzed by binding or segregating in our brain.…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Arda Senocak , Junsik Kim , Tae-Hyun Oh , Hyeonggon Ryu , Dingzeyu Li , In So Kweon

Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be exploited to improve the outputs of models designed for…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-11 Wim Boes , Hugo Van hamme

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable…

Computer Vision and Pattern Recognition · Computer Science 2016-11-11 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to modeling multiple…

Computer Vision and Pattern Recognition · Computer Science 2021-05-05 Yanbei Chen , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…