中文
相关论文

相关论文: Audio Content based Geotagging in Multimedia

200 篇论文

In this paper we describe approaches for discovering acoustic concepts and relations in text. The first major goal is to be able to identify text phrases which contain a notion of audibility and can be termed as a sound or an acoustic…

声音 · 计算机科学 2017-02-14 Anurag Kumar , Bhiksha Raj , Ndapandula Nakashole

The increasing success of audio foundation models across various tasks has led to a growing need for improved interpretability to understand their intricate decision-making processes better. Existing methods primarily focus on explaining…

声音 · 计算机科学 2024-10-11 Alican Akman , Qiyang Sun , Björn W. Schuller

The large size of nowadays' online multimedia databases makes retrieving their content a difficult and time-consuming task. Users of online sound collections typically submit search queries that express a broad intent, often making the…

信息检索 · 计算机科学 2020-06-16 Xavier Favory , Frederic Font , Xavier Serra

This work pioneers the utilization of generative features in enhancing audio understanding. Unlike conventional discriminative features that directly optimize posterior and thus emphasize semantic abstraction while losing fine grained…

声音 · 计算机科学 2025-09-30 Zeyu Xie , Chenxing Li , Xuenan Xu , Mengyue Wu , Wenfu Wang , Ruibo Fu , Meng Yu , Dong Yu , Yuexian Zou

Aerial scene recognition is a fundamental task in remote sensing and has recently received increased interest. While the visual information from overhead images with powerful models and efficient algorithms yields considerable performance…

计算机视觉与模式识别 · 计算机科学 2020-07-17 Di Hu , Xuhong Li , Lichao Mou , Pu Jin , Dong Chen , Liping Jing , Xiaoxiang Zhu , Dejing Dou

Weakly labelled audio tagging aims to predict the classes of sound events within an audio clip, where the onset and offset times of the sound events are not provided. Previous works have used the multiple instance learning (MIL) framework,…

音频与语音处理 · 电气工程与系统科学 2021-02-04 Helin Wang , Yuexian Zou , Wenwu Wang

Spatial audio understanding aims to enable machines to interpret complex auditory scenes, particularly when sound sources move over time. In this work, we study Spatial Audio Question Answering (Spatial AQA) with a focus on movement…

声音 · 计算机科学 2026-02-19 Arvind Krishna Sridhar , Yinyi Guo , Erik Visser

Acoustic event detection is essential for content analysis and description of multimedia recordings. The majority of current literature on the topic learns the detectors through fully-supervised techniques employing strongly labeled data.…

声音 · 计算机科学 2016-07-07 Anurag Kumar , Bhiksha Raj

Large language models have shown a remarkable ability to extract meaning from unstructured data, offering new ways to interpret biomedical signals beyond traditional numerical methods. In this study, we present a matrix factorization…

音频与语音处理 · 电气工程与系统科学 2025-07-15 Yasaman Torabi , Shahram Shirani , James P. Reilly

The problem visual place recognition is commonly used strategy for localization. Most successful appearance based methods typically rely on a large database of views endowed with local or global image descriptors and strive to retrieve the…

计算机视觉与模式识别 · 计算机科学 2016-09-02 Arsalan Mousavian , Jana Kosecka

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

多媒体 · 计算机科学 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

The photogrammetric and reconstructive modeling of cultural heritage sites is mostly focused on visually perceivable aspects, but if their intended purpose is the performance of cultural acts with a sonic emphasis, it is important to…

声音 · 计算机科学 2023-02-14 Dominik Ukolov

In this paper, we focus on the problem of content-based retrieval for audio, which aims to retrieve all semantically similar audio recordings for a given audio clip query. This problem is similar to the problem of query by example of audio,…

声音 · 计算机科学 2018-02-16 Pranay Manocha , Rohan Badlani , Anurag Kumar , Ankit Shah , Benjamin Elizalde , Bhiksha Raj

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

声音 · 计算机科学 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

Perception of auditory events is inherently multimodal relying on both audio and visual cues. A large number of existing multimodal approaches process each modality using modality-specific models and then fuse the embeddings to encode the…

声音 · 计算机科学 2022-07-19 Amir Shirian , Krishna Somandepalli , Victor Sanchez , Tanaya Guha

In this paper we examine the ability of low-level multimodal features to extract movie similarity, in the context of a content-based movie recommendation approach. In particular, we demonstrate the extraction of multimodal representation…

信息检索 · 计算机科学 2019-12-19 Konstantinos Bougiatiotis , Theodore Giannakopoulos

Object-based audio production requires the positional metadata to be defined for each point-source object, including the key elements in the foreground of the sound scene. In many media production use cases, both cameras and microphones are…

音频与语音处理 · 电气工程与系统科学 2024-06-05 Davide Berghi , Philip J. B. Jackson

Conventional audio classification relied on predefined classes, lacking the ability to learn from free-form text. Recent methods unlock learning joint audio-text embeddings from raw audio-text pairs describing audio in natural language.…

多媒体 · 计算机科学 2024-01-11 Ali Vosoughi , Luca Bondi , Ho-Hsiang Wu , Chenliang Xu

As unconventional sources of geo-information, massive imagery and text messages from open platforms and social media form a temporally quasi-seamless, spatially multi-perspective stream, but with unknown and diverse quality. Due to its…

Living in the era of data deluge, we have witnessed a web content explosion, largely due to the massive availability of User-Generated Content (UGC). In this work, we specifically consider the problem of geospatial information extraction…

数据库 · 计算机科学 2013-11-21 Georgios Skoumas , Dieter Pfoser , Anastasios Kyrillidis