English
Related papers

Related papers: ESG-Net: Event-Aware Semantic Guided Network for D…

200 papers

Ensuring the safety of all traffic participants is a prerequisite for bringing intelligent vehicles closer to practical applications. The assistance system should not only achieve high accuracy under normal conditions, but obtain robust…

Computer Vision and Pattern Recognition · Computer Science 2021-12-10 Jiaming Zhang , Kailun Yang , Rainer Stiefelhagen

The use of multiple and semantically correlated sources can provide complementary information to each other that may not be evident when working with individual modalities on their own. In this context, multi-modal models can help producing…

Training deep networks for semantic segmentation requires large amounts of labeled training data, which presents a major challenge in practice, as labeling segmentation masks is a highly labor-intensive process. To address this issue, we…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Lukas Hoyer , Dengxin Dai , Qin Wang , Yuhua Chen , Luc Van Gool

Vision-Language Navigation (VLN) agents often struggle with long-horizon reasoning in unseen environments, particularly when facing ambiguous, coarse-grained instructions. While recent advances use knowledge graph to enhance reasoning, the…

Robotics · Computer Science 2026-03-02 Haoxuan Xu , Tianfu Li , Wenbo Chen , Yi Liu , Xingxing Zuo , Yaoxian Song , Haoang Li

Generic event boundary detection is an important yet challenging task in video understanding, which aims at detecting the moments where humans naturally perceive event boundaries. The main challenge of this task is perceiving various…

Computer Vision and Pattern Recognition · Computer Science 2022-04-04 Jiaqi Tang , Zhaoyang Liu , Chen Qian , Wayne Wu , Limin Wang

This report presents our systems submitted to the audio-only and audio-visual tracks of the DCASE2025 Task 3 Challenge: Stereo Sound Event Localization and Detection (SELD) in Regular Video Content. SELD is a complex task that combines…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-08 Davide Berghi , Philip J. B. Jackson

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Ya Jiang , Qing Wang , Jun Du , Maocheng Hu , Pengfei Hu , Zeyan Liu , Shi Cheng , Zhaoxu Nian , Yuxuan Dong , Mingqi Cai , Xin Fang , Chin-Hui Lee

Sound event detection (SED) has gained increasing attention with its wide application in surveillance, video indexing, etc. Existing models in SED mainly generate frame-level prediction, converting it into a sequence multi-label…

Sound · Computer Science 2021-11-15 Zhirong Ye , Xiangdong Wang , Hong Liu , Yueliang Qian , Rui Tao , Long Yan , Kazushige Ouchi

The goal of weakly supervised video anomaly detection is to learn a detection model using only video-level labeled data. However, prior studies typically divide videos into fixed-length segments without considering the complexity or…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Chen Zhang , Guorong Li , Yuankai Qi , Hanhua Ye , Laiyun Qing , Ming-Hsuan Yang , Qingming Huang

Forecasting Electroncephalography (EEG) signals during cognitive events remains a fundamental challenge in neuroscience and Brain-Computer Interfaces (BCIs), as existing methods struggle to capture both the stochastic nature of neural…

Signal Processing · Electrical Eng. & Systems 2026-03-19 Mehran Shabanpour , Sadaf Khademi , Konstantinos N Plataniotis , Arash Mohammadi

Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA…

Sound · Computer Science 2021-02-12 Yin Cao , Turab Iqbal , Qiuqiang Kong , Fengyan An , Wenwu Wang , Mark D. Plumbley

Document-level Event Causality Identification (DECI) aims to identify causal relations between event pairs in a document. It poses a great challenge of across-sentence reasoning without clear causal indicators. In this paper, we propose a…

Computation and Language · Computer Science 2022-04-18 Meiqi Chen , Yixin Cao , Kunquan Deng , Mukai Li , Kun Wang , Jing Shao , Yan Zhang

The Audio-Visual Event Localization (AVEL) task aims to temporally locate and classify video events that are both audible and visible. Most research in this field assumes a closed-set setting, which restricts these models' ability to handle…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jinxing Zhou , Dan Guo , Ruohao Guo , Yuxin Mao , Jingjing Hu , Yiran Zhong , Xiaojun Chang , Meng Wang

Multimodal entity linking (MEL) aims to utilize multimodal information (usually textual and visual information) to link ambiguous mentions to unambiguous entities in knowledge base. Current methods facing main issues: (1)treating the entire…

Artificial Intelligence · Computer Science 2024-04-11 Shezheng Song , Shasha Li , Shan Zhao , Xiaopeng Li , Chengyu Wang , Jie Yu , Jun Ma , Tianwei Yan , Bin Ji , Xiaoguang Mao

Most existing deep learning-based acoustic scene classification (ASC) approaches directly utilize representations extracted from spectrograms to identify target scenes. However, these approaches pay little attention to the audio events…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-03 Yuanbo Hou , Siyang Song , Chuang Yu , Yuxin Song , Wenwu Wang , Dick Botteldooren

In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Haojian Huang , Kaijing Ma , Jin Chen , Haodong Chen , Zhou Wu , Xianghao Zang , Han Fang , Chao Ban , Hao Sun , Mulin Chen , Zhongjiang He

The Convolutional Neural Networks (CNNs) generate the feature representation of complex objects by collecting hierarchical and different parts of semantic sub-features. These sub-features can usually be distributed in grouped form in the…

Computer Vision and Pattern Recognition · Computer Science 2019-05-28 Xiang Li , Xiaolin Hu , Jian Yang

Event detection (ED), which means identifying event trigger words and classifying event types, is the first and most fundamental step for extracting event knowledge from plain text. Most existing datasets exhibit the following issues that…

Computation and Language · Computer Science 2020-10-09 Xiaozhi Wang , Ziqi Wang , Xu Han , Wangyi Jiang , Rong Han , Zhiyuan Liu , Juanzi Li , Peng Li , Yankai Lin , Jie Zhou

Audio-visual video parsing (AVVP) aims to detect event categories and their temporal boundaries in videos, typically under weak supervision. Existing methods mainly focus on (i) improving temporal modeling using attention-based…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yaru Chen , Faegheh Sardari , Peiliang Zhang , Ruohao Guo , Yang Xiang , Zhenbo Li , Wenwu Wang

Recently, self-supervised learning has proved to be effective to learn representations of events suitable for temporal segmentation in image sequences, where events are understood as sets of temporally adjacent images that are semantically…

Machine Learning · Computer Science 2020-12-11 Mariella Dimiccoli , Herwig Wendt