English
Related papers

Related papers: SwG-former: A Sliding-Window Graph Convolutional N…

200 papers

This report introduces our novel method named STHG for the Audio-Visual Diarization task of the Ego4D Challenge 2023. Our key innovation is that we model all the speakers in a video using a single, unified heterogeneous graph learning…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Kyle Min

Directed graphs are widely used to model asymmetric relationships in real-world systems. However, existing directed graph neural networks often struggle to jointly capture directional semantics and global structural patterns due to their…

Machine Learning · Computer Science 2025-08-20 Jiayu Fang , Zhiqi Shao , S T Boris Choy , Junbin Gao

Spatio-temporal forecasting in various domains, like traffic prediction and weather forecasting, is a challenging endeavor, primarily due to the difficulties in modeling propagation dynamics and capturing high-dimensional interactions among…

Machine Learning · Computer Science 2024-05-29 Xiaobei Zou , Luolin Xiong , Yang Tang , Jürgen Kurths

Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events;…

Sound · Computer Science 2025-08-12 Yuanjian Chen , Yang Xiao , Han Yin , Yadong Guan , Xubo Liu

Cross-modal alignment is an effective approach to improving visual classification. Existing studies typically enforce a one-step mapping that uses deep neural networks to project the visual features to mimic the distribution of textual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zixuan Li , Lei Meng , Guoqing Chao , Wei Wu , Xiaoshuo Yan , Yimeng Yang , Zhuang Qi , Xiangxu Meng

Inspired by the activity-silent and persistent activity mechanisms in human visual perception biology, we design a Unified Static and Dynamic Network (UniSDNet), to learn the semantic association between the video and text/audio queries in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jingjing Hu , Dan Guo , Kun Li , Zhan Si , Xun Yang , Xiaojun Chang , Meng Wang

This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound events and specify their class, furthermore estimate their…

Sound · Computer Science 2025-12-30 Jin Sob Kim , Hyun Joon Park , Wooseok Shin , Sung Won Han

In this paper, we propose the use of spatial and harmonic features in combination with long short term memory (LSTM) recurrent neural network (RNN) for automatic sound event detection (SED) task. Real life sound recordings typically have…

In this technical report, we introduce our solution to human-centric spatio-temporal video grounding task. We propose a concise and effective framework named STVGFormer, which models spatiotemporal visual-linguistic dependencies with a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Zihang Lin , Chaolei Tan , Jian-Fang Hu , Zhi Jin , Tiancai Ye , Wei-Shi Zheng

Adverse weather conditions cause diverse and complex degradation patterns, driving the development of All-in-One (AiO) models. However, recent AiO solutions still struggle to capture diverse degradations, since global filtering methods like…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Yuhwan Jeong , Yunseo Yang , Youngho Yoon , Kuk-Jin Yoon

Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder characterized by complex physiological processes. Previous research has predominantly focused on static cerebral interactions, often neglecting the brain's dynamic nature and…

Machine Learning · Computer Science 2024-09-11 Peng Wang , Xin Wen , Ruochen Cao , Chengxin Gao , Yanrong Hao , Rui Cao

Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-22 Hengyi Hong , Qing Wang , Jun Du , Ruoyu Wei , Mingqi Cai , Xin Fang

Video scene graph generation (VidSGG) aims to identify objects in visual scenes and infer their relationships for a given video. It requires not only a comprehensive understanding of each object scattered on the whole scene but also a deep…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 Tao Pu , Tianshui Chen , Hefeng Wu , Yongyi Lu , Liang Lin

As deeper and more complex models are developed for the task of sound event localization and detection (SELD), the demand for annotated spatial audio data continues to increase. Annotating field recordings with 360$^{\circ}$ video takes…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-08 Christopher Ick , Brian McFee

In this report, we propose three novel methods for developing a sound event detection (SED) model for the DCASE 2024 Challenge Task 4. First, we propose an auxiliary decoder attached to the final convolutional block to improve feature…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-25 Sang Won Son , Jongyeon Park , Hong Kook Kim , Sulaiman Vesal , Jeong Eun Lim

We explore on various attention methods on frequency and channel dimensions for sound event detection (SED) in order to enhance performance with minimal increase in computational cost while leveraging domain knowledge to address the…

Sound · Computer Science 2023-08-30 Hyeonuk Nam , Seong-Hu Kim , Deokki Min , Yong-Hwa Park

Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semantic constraints on the final outputs, lacking cross-modal…

Multimedia · Computer Science 2025-10-16 Huilai Li , Yonghao Dang , Ying Xing , Yiming Wang , Jianqin Yin

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting…

Sound · Computer Science 2025-10-28 Pengfei Cai , Yan Song , Qing Gu , Nan Jiang , Haoyu Song , Ian McLoughlin

Polyphonic sound event localization and detection (SELD) has many practical applications in acoustic sensing and monitoring. However, the development of real-time SELD has been limited by the demanding computational requirement of most…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-07 Thi Ngoc Tho Nguyen , Douglas L. Jones , Karn N. Watcharasupat , Huy Phan , Woon-Seng Gan

Scene Graph Generation (SGG) serves a comprehensive representation of the images for human understanding as well as visual understanding tasks. Due to the long tail bias problem of the object and predicate labels in the available annotated…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Anh Duc Bui , Soyeon Caren Han , Josiah Poon