中文
相关论文

相关论文: City-Identification of Flickr Videos Using Semanti…

200 篇论文

In this paper we propose methods to extract geographically relevant information in a multimedia recording using its audio. Our method primarily is based on the fact that urban acoustic environment consists of a variety of sounds. Hence,…

声音 · 计算机科学 2016-11-14 Anurag Kumar , Benjamin Elizalde , Bhiksha Raj

Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner, and by design excludes temporal information present in videos. While it proves to be effective for…

声音 · 计算机科学 2023-04-18 Rajsuryan Singh , Pablo Zinemanas , Xavier Serra , Juan Pablo Bello , Magdalena Fuentes

The objective of this work is to localize sound sources that are visible in a video without using manual annotations. Our key technical contribution is to show that, by training the network to explicitly discriminate challenging image…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Honglie Chen , Weidi Xie , Triantafyllos Afouras , Arsha Nagrani , Andrea Vedaldi , Andrew Zisserman

Unsupervised audio-visual source localization aims at localizing visible sound sources in a video without relying on ground-truth localization for training. Previous works often seek high audio-visual similarities for likely positive…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Shentong Mo , Pedro Morgado

Environmental soundscapes convey substantial ecological and social information regarding urban environments; however, their potential remains largely untapped in large-scale geographic analysis. In this study, we investigate the extent to…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Pengyu Chen , Xiao Huang , Teng Fei , Sicheng Wang

The majority of sound scene analysis work focuses on one of two clearly defined tasks: acoustic scene classification or sound event detection. Whilst this separation of tasks is useful for problem definition, they inherently ignore some…

音频与语音处理 · 电气工程与系统科学 2019-07-30 Helen L. Bear , Toni Heittola , Annamaria Mesaros , Emmanouil Benetos , Tuomas Virtanen

Learning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to…

计算机视觉与模式识别 · 计算机科学 2022-11-08 Dennis Fedorishin , Deen Dayal Mohan , Bhavin Jawade , Srirangaraj Setlur , Venu Govindaraju

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extremely simple approach to generating (weak) speech…

多媒体 · 计算机科学 2017-06-02 Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , Ian Sturdy

Having knowledge of the environmental context of the user i.e. the knowledge of the users' indoor location and the semantics of their environment, can facilitate the development of many of location-aware applications. In this paper, we…

声音 · 计算机科学 2018-04-03 Muhammad A. Shah , Bhiksha Raj , Khaled A. Harras

Can we determine someone's geographic location purely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? We tackle the challenge of global-scale audio geolocation, formalize the…

声音 · 计算机科学 2025-07-23 Mustafa Chasmai , Wuao Liu , Subhransu Maji , Grant Van Horn

Acoustic scene recordings are often collected from a diverse range of cities. Most existing acoustic scene classification (ASC) approaches focus on identifying common acoustic scene patterns across cities to enhance generalization. However,…

声音 · 计算机科学 2025-06-16 Yiqiang Cai , Yizhou Tan , Shengchen Li , Xi Shao , Mark D. Plumbley

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu

We present SONYC-UST-V2, a dataset for urban sound tagging with spatiotemporal information. This dataset is aimed for the development and evaluation of machine listening systems for real-world urban noise monitoring. While datasets of urban…

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

声音 · 计算机科学 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Audio tagging aims to label sound events appearing in an audio recording. In this paper, we propose region-specific audio tagging, a new task which labels sound events in a given region for spatial audio recorded by a microphone array. The…

音频与语音处理 · 电气工程与系统科学 2025-09-12 Jinzheng Zhao , Yong Xu , Haohe Liu , Davide Berghi , Xinyuan Qian , Qiuqiang Kong , Junqi Zhao , Mark D. Plumbley , Wenwu Wang

Sound scene geotagging is a new topic of research which has evolved from acoustic scene classification. It is motivated by the idea of audio surveillance. Not content with only describing a scene in a recording, a machine which can locate…

音频与语音处理 · 电气工程与系统科学 2021-10-12 Helen L. Bear , Veronica Morfi , Emmanouil Benetos

During the performance of sound source localization which uses both visual and aural information, it presently remains unclear how much either image or sound modalities contribute to the result, i.e. do we need both image and sound for…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Takashi Oya , Shohei Iwase , Ryota Natsume , Takahiro Itazuri , Shugo Yamaguchi , Shigeo Morishima

The identification of device brands and models plays a pivotal role in the realm of multimedia forensic applications. This paper presents a framework capable of identifying devices using audio, visual content, or a fusion of them. The…

机器学习 · 计算机科学 2024-06-27 Ioannis Tsingalis , Christos Korgialas , Constantine Kotropoulos

The objective of the sound source localization task is to enable machines to detect the location of sound-making objects within a visual scene. While the audio modality provides spatial cues to locate the sound source, existing approaches…

多媒体 · 计算机科学 2023-08-21 Sung Jin Um , Dongjin Kim , Jung Uk Kim
‹ 上一页 1 2 3 10 下一页 ›