English
Related papers

Related papers: TMac: Temporal Multi-Modal Graph Learning for Acou…

200 papers

With the development of media and networking technologies, multimedia applications ranging from feature presentation in a cinema setting to video on demand to interactive video conferencing are in great demand. Good synchronization between…

Computer Vision and Pattern Recognition · Computer Science 2018-12-17 Naji Khosravan , Shervin Ardeshir , Rohit Puri

While end-to-end systems are becoming popular in auditory signal processing including automatic music tagging, models using raw audio as input needs a large amount of data and computational resources without domain knowledge. Inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-29 Yinghao Ma , Richard M. Stern

This work presents MAD (Multimodal Affection Dataset), a multimodal emotion dataset designed for affective computing and neurophysiological modeling. MAD is built upon synchronous collection of diverse physiological signals (EEG, ECG, EOG,…

Signal Processing · Electrical Eng. & Systems 2026-03-09 Shengwei Guo , Yunqing Qiao , Wenzhan Zhang , Bo Liu , Yong Wang , Guobing Sun

Temporal Video Grounding (TVG) aims to localize the temporal boundary of a specific segment in an untrimmed video based on a given language query. Since datasets in this domain are often gathered from limited video scenes, models tend to…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Haifeng Huang , Yang Zhao , Zehan Wang , Yan Xia , Zhou Zhao

Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introduce Timestamped Audio Captioner (TAC), a model that produces…

Temporal Knowledge Graph Forecasting (TKGF) aims to predict future events based on the observed events in history. Recently, Large Language Models (LLMs) have exhibited remarkable capabilities, generating significant research interest in…

Information Retrieval · Computer Science 2025-01-22 He Chang , Jie Wu , Zhulin Tao , Yunshan Ma , Xianglin Huang , Tat-Seng Chua

This paper focuses on the weakly-supervised audio-visual video parsing task, which aims to recognize all events belonging to each modality and localize their temporal boundaries. This task is challenging because only overall labels…

Computer Vision and Pattern Recognition · Computer Science 2022-08-02 Haoyue Cheng , Zhaoyang Liu , Hang Zhou , Chen Qian , Wayne Wu , Limin Wang

Temporal knowledge graphs (TKGs) inherently reflect the transient nature of real-world knowledge, as opposed to static knowledge graphs. Naturally, automatic TKG completion has drawn much research interests for a more realistic modeling of…

Machine Learning · Computer Science 2020-12-22 Jaehun Jung , Jinhong Jung , U Kang

Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS). However, this setting might be unrealistic as different classes of actions…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Jing Tan , Xiaotong Zhao , Xintian Shi , Bin Kang , Limin Wang

Effectively leveraging multimodal data such as various images, laboratory tests and clinical information is gaining traction in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on…

Image and Video Processing · Electrical Eng. & Systems 2023-11-28 Yingying Fang , Shuang Wu , Sheng Zhang , Chaoyan Huang , Tieyong Zeng , Xiaodan Xing , Simon Walsh , Guang Yang

Current optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel in challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Qianang Zhou , Junhui Hou , Meiyi Yang , Yongjian Deng , Youfu Li , Junlin Xiong

Learning human motion based on a time-dependent input signal presents a challenging yet impactful task with various applications. The goal of this task is to generate or estimate human movement that consistently reflects the temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Quang Nguyen , Tri Le , Baoru Huang , Minh Nhat Vu , Ngan Le , Thieu Vo , Anh Nguyen

Audio-visual representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel multimodal framework that instantiates multiple instance…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Sanjeel Parekh , Slim Essid , Alexey Ozerov , Ngoc Q. K. Duong , Patrick Pérez , Gaël Richard

Deep learning models have enjoyed great success for image related computer vision tasks like image classification and object detection. For video related tasks like human action recognition, however, the advancements are not as significant…

Computer Vision and Pattern Recognition · Computer Science 2018-09-12 Xiaolin Song , Cuiling Lan , Wenjun Zeng , Junliang Xing , Jingyu Yang , Xiaoyan Sun

The widespread availability of complex time series data in various domains such as environmental science, epidemiology, and economics demands robust causal discovery methods that can identify intricate contemporaneous and lagged…

Machine Learning · Computer Science 2026-05-12 Omar Faruque , Sahara Ali , Xue Zheng , Jianwu Wang

Most existing multimodal trackers adopt uniform fusion strategies, overlooking the inherent differences between modalities. Moreover, they propagate temporal information through mixed tokens, leading to entangled and less discriminative…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Shilei Wang , Pujian Lai , Dong Gao , Jifeng Ning , Gong Cheng

Multi-modal learning has shown exceptional performance in various tasks, especially in medical applications, where it integrates diverse medical information for comprehensive diagnostic evidence. However, there still are several challenges…

Machine Learning · Computer Science 2024-11-19 Lin Fan , Yafei Ou , Cenyang Zheng , Pengyu Dai , Tamotsu Kamishima , Masayuki Ikebe , Kenji Suzuki , Xun Gong

Despite the recent success of machine learning algorithms, most models face drawbacks when considering more complex tasks requiring interaction between different sources, such as multimodal input data and logical time sequences. On the…

Sound · Computer Science 2023-02-01 Leandro A. Passos , João Paulo Papa , Amir Hussain , Ahsan Adeel

Dynamic graph learning is essential for applications involving temporal networks and requires effective modeling of temporal relationships. Seminal attention-based models like TGAT and DyGFormer rely on sinusoidal time encoders to capture…

Machine Learning · Computer Science 2025-08-05 Hsing-Huan Chung , Shravan Chaudhari , Xing Han , Yoav Wald , Suchi Saria , Joydeep Ghosh

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Haytham M. Fayek , Anurag Kumar