中文
相关论文

相关论文: Localizing Visual Sounds the Hard Way

200 篇论文

In this paper, we introduce a simple and standalone manual annotation tool for images, audio and video: the VGG Image Annotator (VIA). This is a light weight, standalone and offline software package that does not require any installation or…

计算机视觉与模式识别 · 计算机科学 2019-08-12 Abhishek Dutta , Andrew Zisserman

We explore means to advance source camera identification based on sensor noise in a data-driven framework. Our focus is on improving the sensor pattern noise (SPN) extraction from a single image at test time. Where existing works suppress…

计算机视觉与模式识别 · 计算机科学 2020-02-10 Matthias Kirchner , Cameron Johnson

The text detection and localization plays a major role in video analysis and understanding. The scene text embedded in video consist of high-level semantics and hence contributes significantly to visual content analysis and retrieval. This…

计算机视觉与模式识别 · 计算机科学 2015-02-24 B. H. Shekar , Smitha M. L.

Annotating images with pixel-wise labels is a time-consuming and costly process. Recently, DatasetGAN showcased a promising alternative - to synthesize a large labeled dataset via a generative adversarial network (GAN) by exploiting a small…

计算机视觉与模式识别 · 计算机科学 2022-01-14 Daiqing Li , Huan Ling , Seung Wook Kim , Karsten Kreis , Adela Barriuso , Sanja Fidler , Antonio Torralba

Surgical tool localization is an essential task for the automatic analysis of endoscopic videos. In the literature, existing methods for tool localization, tracking and segmentation require training data that is fully annotated, thereby…

计算机视觉与模式识别 · 计算机科学 2018-07-19 Armine Vardazaryan , Didier Mutter , Jacques Marescaux , Nicolas Padoy

Automatically describing audio-visual content with texts, namely video captioning, has received significant attention due to its potential applications across diverse fields. Deep neural networks are the dominant methods, offering…

音频与语音处理 · 电气工程与系统科学 2023-06-19 Özkan Çaylı , Xubo Liu , Volkan Kılıç , Wenwu Wang

A long-standing goal in the field of sensory substitution is to enable sound perception for deaf and hard of hearing (DHH) people by visualizing audio content. Different from existing models that translate to hand sign language, between…

人机交互 · 计算机科学 2023-02-15 Chunjin Song , Yuchi Zhang , Willis Peng , Parmis Mohaghegh , Bastian Wandt , Helge Rhodin

Video transcript summarization is a fundamental task for video understanding. Conventional approaches for transcript summarization are usually built upon the summarization data for written language such as news articles, while the domain…

计算与语言 · 计算机科学 2021-07-16 Tengchao Lv , Lei Cui , Momcilo Vasilijevic , Furu Wei

Over the past few years, there has been a great deal of research on navigation tasks in indoor environments using deep reinforcement learning agents. Most of these tasks use only visual information in the form of first-person images to…

计算机视觉与模式识别 · 计算机科学 2023-08-02 Haru Kondoh , Asako Kanezaki

Video object segmentation (VOS) has made significant progress with the rise of deep learning. However, there still exist some thorny problems, for example, similar objects are easily confused and tiny objects are difficult to be found. To…

计算机视觉与模式识别 · 计算机科学 2022-06-22 Wangwang Yang , Jinming Su , Yiting Duan , Tingyi Guo , Junfeng Luo

Sound source tracking is commonly performed using classical array-processing algorithms, while machine-learning approaches typically rely on precise source position labels that are expensive or impractical to obtain. This paper introduces a…

音频与语音处理 · 电气工程与系统科学 2026-02-12 Luan Vinícius Fiorio , Ivana Nikoloska , Bruno Defraene , Alex Young , Johan David , Ronald M. Aarts

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Vladimir Iashin , Weidi Xie , Esa Rahtu , Andrew Zisserman

Video semantic segmentation (VSS) is beneficial for dealing with dynamic scenes due to the continuous property of the real-world environment. On the one hand, some methods alleviate the predicted inconsistent problem between continuous…

计算机视觉与模式识别 · 计算机科学 2023-01-02 Yuhang Zhang , Shishun Tian , Muxin Liao , Zhengyu Zhang , Wenbin Zou , Chen Xu

Recent studies have advocated the detection of fake videos as a one-class detection task, predicated on the hypothesis that the consistency between audio and visual modalities of genuine data is more significant than that of fake data. This…

声音 · 计算机科学 2024-06-13 Xiaolou Li , Zehua Liu , Chen Chen , Lantian Li , Li Guo , Dong Wang

The text detection and localization is important for video analysis and understanding. The scene text in video contains semantic information and thus can contribute significantly to video retrieval and understanding. However, most of the…

计算机视觉与模式识别 · 计算机科学 2015-02-25 B. H. Shekar , Smitha M. L. , P. Shivakumara

Deep learning has made significant strides in video understanding tasks, but the computation required to classify lengthy and massive videos using clip-level video classifiers remains impractical and prohibitively expensive. To address this…

计算机视觉与模式识别 · 计算机科学 2023-08-21 Muhammad Adi Nugroho , Sangmin Woo , Sumin Lee , Changick Kim

The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription or OCR-based…

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

Sound sources localization using multichannel signal processing has been a subject of active research for decades. In recent years, the use of deep learning in audio signal processing has allowed to drastically improve performances for…

音频与语音处理 · 电气工程与系统科学 2021-06-16 Hadrien Pujol , Éric Bavu , Alexandre Garcia

In our work, we examine, for the first time, the possibility of fast and efficient source localization directly from the uvobservations, omitting the recovering of the dirty or clean images. We propose a deep neural network-based framework…

天体物理仪器与方法 · 物理学 2023-06-21 O. Taran , O. Bait , M. Dessauges-Zavadsky , T. Holotyak , D. Schaerer , S. Voloshynovskiy
‹ 上一页 1 8 9 10 下一页 ›