中文
相关论文

相关论文: Exploring Fine-Grained Audiovisual Categorization …

200 篇论文

We explore a new task for audio-visual-language modeling called fine-grained audible video description (FAVD). It aims to provide detailed textual descriptions for the given audible videos, including the appearance and spatial locations of…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Xuyang Shen , Dong Li , Jinxing Zhou , Zhen Qin , Bowen He , Xiaodong Han , Aixuan Li , Yuchao Dai , Lingpeng Kong , Meng Wang , Yu Qiao , Yiran Zhong

Manual labeling of animal images remains a significant bottleneck in ecological research, limiting the scale and efficiency of biodiversity monitoring efforts. This study investigates whether state-of-the-art Vision Transformer (ViT)…

计算机视觉与模式识别 · 计算机科学 2026-02-05 Hugo Markoff , Stefan Hein Bengtson , Michael Ørsted

In this paper, we propose two techniques, namely joint modeling and data augmentation, to improve system performances for audio-visual scene classification (AVSC). We employ pre-trained networks trained only on image data sets to extract…

Image fusion aims to synthesize a single high-quality image from a pair of inputs captured under challenging conditions, such as differing exposure levels or focal depths. A core challenge lies in effectively handling disparities in dynamic…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Mingwei Tang , Jiahao Nie , Guang Yang , Ziqing Cui , Jie Li

We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attention-based feedforward structure in such a way that allows the…

音频与语音处理 · 电气工程与系统科学 2019-12-06 Wim Boes , Hugo Van hamme

This study aims to explore different pre-trained models offered in the Torchvision package which is available in the PyTorch library. And investigate their effectiveness on fine-grained images classification. Transfer Learning is an…

计算机视觉与模式识别 · 计算机科学 2021-10-15 Feras Albardi , H M Dipu Kabir , Md Mahbub Islam Bhuiyan , Parham M. Kebria , Abbas Khosravi , Saeid Nahavandi

Prior work has studied different visual modalities in isolation and developed separate architectures for recognition of images, videos, and 3D data. Instead, in this paper, we propose a single model which excels at classifying images,…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Rohit Girdhar , Mannat Singh , Nikhila Ravi , Laurens van der Maaten , Armand Joulin , Ishan Misra

We introduce the new Birds-to-Words dataset of 41k sentences describing fine-grained differences between photographs of birds. The language collected is highly detailed, while remaining understandable to the everyday observer (e.g.,…

计算与语言 · 计算机科学 2019-11-15 Maxwell Forbes , Christine Kaeser-Chen , Piyush Sharma , Serge Belongie

The current biodiversity loss crisis makes animal monitoring a relevant field of study. In light of this, data collected through monitoring can provide essential insights, and information for decision-making aimed at preserving global…

In the following paper, we present and discuss challenging applications for fine-grained visual classification (FGVC): biodiversity and species analysis. We not only give details about two challenging new datasets suitable for computer…

计算机视觉与模式识别 · 计算机科学 2015-07-06 Erik Rodner , Marcel Simon , Gunnar Brehm , Stephanie Pietsch , J. Wolfgang Wägele , Joachim Denzler

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos,…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Lin Zhang , Zefan Cai , Yufan Zhou , Shentong Mo , Jinhong Lin , Cheng-En Wu , Yibing Wei , Yijing Zhang , Ruiyi Zhang , Wen Xiao , Tong Sun , Junjie Hu , Pedro Morgado

Visible and infrared image fusion (VIF) has gained significant attention in recent years due to its wide application in tasks such as scene segmentation and object detection. VIF methods can be broadly classified into traditional VIF…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Zixian Zhao , Xingchen Zhang

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

This study introduces a novel multimodal food recognition framework that effectively combines visual and textual modalities to enhance classification accuracy and robustness. The proposed approach employs a dynamic multimodal fusion…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Prateek Mittal , Puneet Goyal , Joohi Chauhan

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, have demonstrated strong alignment between images and textual…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Risa Shinoda , Kaede Shiohara , Nakamasa Inoue , Kuniaki Saito , Hiroaki Santo , Fumio Okura

Video-text retrieval has seen significant advancements, yet the ability of models to discern subtle differences in captions still requires verification. In this paper, we introduce a new approach for fine-grained evaluation. Our approach…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Aozhu Chen , Hazel Doughty , Xirong Li , Cees G. M. Snoek

We study the merit of transfer learning for two sound recognition problems, i.e., audio tagging and sound event detection. Employing feature fusion, we adapt a baseline system utilizing only spectral acoustic inputs to also make use of…

音频与语音处理 · 电气工程与系统科学 2022-09-27 Wim Boes , Hugo Van hamme

Fine-Grained Visual Classification(FGVC) is the task that requires recognizing the objects belonging to multiple subordinate categories of a super-category. Recent state-of-the-art methods usually design sophisticated learning pipelines to…

计算机视觉与模式识别 · 计算机科学 2022-03-08 Qishuai Diao , Yi Jiang , Bin Wen , Jia Sun , Zehuan Yuan

Social media platforms enable the propagation of hateful content across different modalities such as textual, auditory, and visual, necessitating effective detection methods. While recent approaches have shown promise in handling individual…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Girish A. Koushik , Diptesh Kanojia , Helen Treharne

Audio-Visual Segmentation (AVS) aims to identify and segment sound-producing objects in videos by leveraging both visual and audio modalities. It has emerged as a significant research area in multimodal perception, enabling fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Jia Li , Yapeng Tian