English
Related papers

Related papers: Audio Content based Geotagging in Multimedia

200 papers

Sound waves cause small vibrations in nearby objects. A few techniques exist in the literature that can extract sound from video. In this paper we study local vibration patterns at different image locations. We show that different locations…

Computer Vision and Pattern Recognition · Computer Science 2019-07-16 Mohammad Amin Shabani , Laleh Samadfam , Mohammad Amin Sadeghi

Geotagged data can be used to describe regions in the world and discover local themes. However, not all data produced within a region is necessarily specifically descriptive of that area. To surface the content that is characteristic for a…

Machine Learning · Statistics 2015-03-13 Mohamed Kafsi , Henriette Cramer , Bart Thomee , David A. Shamma

Human auditory perception is shaped by moving sound sources in 3D space, yet prior work in generative sound modelling has largely been restricted to mono signals or static spatial audio. In this work, we introduce a framework for generating…

Sound · Computer Science 2025-09-29 Yunyi Liu , Shaofan Yang , Kai Li , Xu Li

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

Graphics · Computer Science 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label…

Sound · Computer Science 2025-07-17 Jaesung Huh , Jacob Chalk , Evangelos Kazakos , Dima Damen , Andrew Zisserman

Most universal sound extraction algorithms focus on isolating a target sound event from single-channel audio mixtures. However, the real world is three-dimensional, and binaural audio, which mimics human hearing, can capture richer spatial…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Zexu Pan , Shengkui Zhao , Yukun Ma , Haoxu Wang , Yiheng Jiang , Biao Tian , Bin Ma

A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the problem of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Chuang Gan , Yiwei Zhang , Jiajun Wu , Boqing Gong , Joshua B. Tenenbaum

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba

We here summarize our experience running a challenge with open data for musical genre recognition. Those notes motivate the task and the challenge design, show some statistics about the submissions, and present the results.

Sound · Computer Science 2018-03-15 Michaël Defferrard , Sharada P. Mohanty , Sean F. Carroll , Marcel Salathé

This paper introduces a new paradigm for sound source lo-calization referred to as virtual acoustic space traveling (VAST) and presents a first dataset designed for this purpose. Existing sound source localization methods are either based…

Sound · Computer Science 2016-12-20 Clément Gaultier , Saurabh Kataria , Antoine Deleforge

Due to the extensive use of information technology and the recent developments in multimedia systems, the amount of multimedia data available to users has increased exponentially. Video is an example of multimedia data as it contains…

Information Retrieval · Computer Science 2014-01-03 Avinash N Bhute , B B Meshram

This study presents a cascaded architecture for extractive summarization of multimedia content via audio-to-text alignment. The proposed framework addresses the challenge of extracting key insights from multimedia sources like YouTube…

Information Retrieval · Computer Science 2025-04-10 Tanzir Hossain , Ar-Rafi Islam , Md. Sabbir Hossain , Annajiat Alim Rasel

Content-based information retrieval is based on the information contained in documents rather than using metadata such as keywords. Most information retrieval methods are either based on text or image. In this paper, we investigate the…

Computation and Language · Computer Science 2020-10-02 Golsa Tahmasebzadeh , Sherzod Hakimov , Eric Müller-Budack , Ralph Ewerth

In this paper, we propose to infer music genre embeddings from audio datasets carrying semantic information about genres. We show that such embeddings can be used for disambiguating genre tags (identification of different labels for the…

Information Retrieval · Computer Science 2018-09-20 Romain Hennequin , Jimena Royo-Letelier , Manuel Moussallam

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

Optoacoustic image formation is conventionally based upon ultrasound time-of-flight readings from multiple detection positions. Herein, we exploit acoustic scattering to physically encode the position of optical absorbers in the acquired…

Biological Physics · Physics 2019-10-30 Xose Luis Dean-Ben , Ali Ozbek , Hernan Lopez-Schier , Daniel Razansky

During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not…

Sound · Computer Science 2022-07-22 Wei Sun , Mei Wang , Lili Qiu

Maps are crucial in conveying geospatial data in diverse contexts such as news and scientific reports. This research, utilizing thematic maps, probes deeper into the underexplored intersection of text framing and map types in influencing…

Human-Computer Interaction · Computer Science 2024-03-14 Arlen Fan , Fan Lei , Michelle Mancenido , Alan MacEachren , Ross Maciejewski

The localization of sound sources by the human brain is computationally simulated from a neurobiological perspective. The simulation includes the neural representation of temporal differences in acoustic signals between the ipsilateral and…

Neurons and Cognition · Quantitative Biology 2008-10-31 Nikesh S. Dattani

Food containers, drinking glasses and cups handled by a person generate sounds that vary with the type and amount of their content. In this paper, we propose a new model for sound-based classification of the type and amount of content in a…

Sound · Computer Science 2021-06-10 Santiago Donaher , Alessio Xompero , Andrea Cavallaro
‹ Prev 1 4 5 6 7 8 10 Next ›