中文
相关论文

相关论文: UniAV: Unified Audio-Visual Perception for Multi-T…

200 篇论文

Open-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabilities. Closed-vocab TAD is typically confined to localizing…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Trung Thanh Nguyen , Yasutomo Kawanishi , Takahiro Komamizu , Ichiro Ide

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Joseph Heyward , Nikhil Parthasarathy , Tyler Zhu , Aravindh Mahendran , João Carreira , Dima Damen , Andrew Zisserman , Viorica Pătrăucean

Weakly Labelled learning has garnered lot of attention in recent years due to its potential to scale Sound Event Detection (SED) and is formulated as Multiple Instance Learning (MIL) problem. This paper proposes a Multi-Task Learning (MTL)…

音频与语音处理 · 电气工程与系统科学 2020-11-02 Soham Deshmukh , Bhiksha Raj , Rita Singh

Sound Event Localization and Detection (SELD) is crucial in spatial audio processing, enabling systems to detect sound events and estimate their 3D directions. Existing SELD methods use single- or dual-branch architectures: single-branch…

声音 · 计算机科学 2025-07-31 Hogeon Yu

Sound event localization and detection (SELD) is a joint task of sound event detection and direction-of-arrival estimation. In DCASE 2022 Task 3, types of data transform from computationally generated spatial recordings to recordings of…

音频与语音处理 · 电气工程与系统科学 2022-09-12 Jinbo Hu , Yin Cao , Ming Wu , Qiuqiang Kong , Feiran Yang , Mark D. Plumbley , Jun Yang

Animal visual perception is an important technique for automatically monitoring animal health, understanding animal behaviors, and assisting animal-related research. However, it is challenging to design a deep learning-based perception…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Meiqi Sun , Zhonghan Zhao , Wenhao Chai , Hanjun Luo , Shidong Cao , Yanting Zhang , Jenq-Neng Hwang , Gaoang Wang

This paper addresses the task of unsupervised video multi-object segmentation. Current approaches follow a two-stage paradigm: 1) detect object proposals using pre-trained Mask R-CNN, and 2) conduct generic feature matching for temporal…

计算机视觉与模式识别 · 计算机科学 2021-04-13 Tianfei Zhou , Jianwu Li , Xueyi Li , Ling Shao

Video action recognition is a challenging but important task for understanding and discovering what the video does. However, acquiring annotations for a video is costly, and semi-supervised learning (SSL) has been studied to improve…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Seokun Kang , Taehwan Kim

The forensic investigation of a terrorist attack poses a significant challenge to the investigative authorities, as often several thousand hours of video footage must be viewed. Large scale Video Analytic Platforms (VAP) assist law…

多媒体 · 计算机科学 2020-04-03 Alexander Schindler , Andrew Lindley , Anahid Jalali , Martin Boyer , Sergiu Gordea , Ross King

The Audio-Visual Video Parsing task aims to identify and temporally localize the events that occur in either or both the audio and visual streams of audible videos. It often performs in a weakly-supervised manner, where only video event…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Jinxing Zhou , Dan Guo , Yiran Zhong , Meng Wang

Advanced driver assistance systems require a comprehensive understanding of the driver's mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Wenzhuo Liu , Wenshuo Wang , Yicheng Qiao , Qiannan Guo , Jiayin Zhu , Pengfei Li , Zilong Chen , Huiming Yang , Zhiwei Li , Lening Wang , Tiao Tan , Huaping Liu

Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Zixuan Ye , Xuanhua He , Quande Liu , Qiulin Wang , Xintao Wang , Pengfei Wan , Di Zhang , Kun Gai , Qifeng Chen , Wenhan Luo

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Jiabin Luo , Junhui Lin , Zeyu Zhang , Biao Wu , Meng Fang , Ling Chen , Hao Tang

In audio-visual navigation (AVN), an intelligent agent needs to navigate to a constantly sound-making object in complex 3D environments based on its audio and visual perceptions. While existing methods attempt to improve the navigation…

声音 · 计算机科学 2022-06-02 Shunqi Mao , Chaoyi Zhang , Heng Wang , Weidong Cai

Sound event detection (SED) entails identifying the type of sound and estimating its temporal boundaries from acoustic signals. These events are uniquely characterized by their spatio-temporal features, which are determined by the way they…

音频与语音处理 · 电气工程与系统科学 2023-05-19 Tanmay Khandelwal , Rohan Kumar Das

Scene understanding is crucial for autonomous systems which intend to operate in the real world. Single task vision networks extract information only based on some aspects of the scene. In multi-task learning (MTL), on the other hand, these…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Naresh Kumar Gurulingan , Elahe Arani , Bahram Zonooz

The human brain can easily focus on one speaker and suppress others in scenarios such as a cocktail party. Recently, researchers found that auditory attention can be decoded from the electroencephalogram (EEG) data. However, most existing…

声音 · 计算机科学 2023-08-09 Xiaoyu Chen , Changde Du , Qiongyi Zhou , Huiguang He

Developing Audio-Visual Large Language Models (AV-LLMs) for unified scene understanding is pivotal in multimodal intelligence. While instruction tuning enables pre-trained models with multi-task abilities, we observe that conventional…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Dongnuan Cai , Henghui Du , Chang Zhou , Xi Chen , Dan Guo , Hongyuan Zhang , Xuelong Li , Di Hu

Acoustic event detection and scene classification are major research tasks in environmental sound analysis, and many methods based on neural networks have been proposed. Conventional methods have addressed these tasks separately; however,…

Audio-visual learning, aimed at exploiting the relationship between audio and visual modalities, has drawn considerable attention since deep learning started to be used successfully. Researchers tend to leverage these two modalities either…

计算机视觉与模式识别 · 计算机科学 2020-01-15 Hao Zhu , Mandi Luo , Rui Wang , Aihua Zheng , Ran He