English
Related papers

Related papers: TMac: Temporal Multi-Modal Graph Learning for Acou…

200 papers

Graph machine learning has made significant strides in recent years, yet the integration of visual information with graph structure and its potential for improving performance in downstream tasks remains an underexplored area. To address…

Machine Learning · Computer Science 2025-04-01 Jing Zhu , Yuhang Zhou , Shengyi Qian , Zhongmou He , Tong Zhao , Neil Shah , Danai Koutra

We tackle the task of environmental event classification by drawing inspiration from the transformer neural network architecture used in machine translation. We modify this attention-based feedforward structure in such a way that allows the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-06 Wim Boes , Hugo Van hamme

Session-based recommendation plays a central role in a wide spectrum of online applications, ranging from e-commerce to online advertising services. However, the majority of existing session-based recommendation techniques (e.g.,…

Information Retrieval · Computer Science 2021-10-11 Chao Huang , Jiahui Chen , Lianghao Xia , Yong Xu , Peng Dai , Yanqing Chen , Liefeng Bo , Jiashu Zhao , Jimmy Xiangji Huang

Multi-modal time series analysis has recently emerged as a prominent research area in data mining, driven by the increasing availability of diverse data modalities, such as text, images, and structured tabular data from real-world sources.…

Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-driven methods cannot…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Fangyi Zhu , Jenq-Neng Hwang , Zhanyu Ma , Guang Chen , Jun Guo

Multi-modal object tracking has attracted considerable attention by integrating multiple complementary inputs (e.g., thermal, depth, and event data) to achieve outstanding performance. Although current general-purpose multi-modal trackers…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Qihua Liang , Liang Chen , Yaozong Zheng , Jian Nong , Zhiyi Mo , Bineng Zhong

Multimodal Sentiment Analysis (MSA) endeavors to understand human sentiment by leveraging language, visual, and acoustic modalities. Despite the remarkable performance exhibited by previous MSA approaches, the presence of inherent…

Multimedia · Computer Science 2025-05-09 Weize Quan , Yunfei Feng , Ming Zhou , Yunzhen Zhao , Tong Wang , Dong-Ming Yan

This paper explores the development of a multimodal sentiment analysis model that integrates text, audio, and visual data to enhance sentiment classification. The goal is to improve emotion detection by capturing the complex interactions…

Computation and Language · Computer Science 2025-01-15 Hui Lee , Singh Suniljit , Yong Siang Ong

Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval.…

Sound · Computer Science 2022-10-31 Kleanthis Avramidis , Shanti Stewart , Shrikanth Narayanan

While tabular machine learning has achieved remarkable success, temporal distribution shifts pose significant challenges in real-world deployment, as the relationships between features and labels continuously evolve. Static models assume…

Machine Learning · Computer Science 2025-12-04 Hao-Run Cai , Han-Jia Ye

Multimodal recommendation systems have attracted increasing attention for their improved performance by leveraging items' multimodal information. Prior methods often build modality-specific item-item semantic graphs from raw modality…

Information Retrieval · Computer Science 2025-08-11 Xiaoxiong Zhang , Xin Zhou , Zhiwei Zeng , Dusit Niyato , Zhiqi Shen

Controllable text-to-audio generation aims to synthesize audio from textual descriptions while satisfying user-specified constraints, including event types, temporal sequences, and onset and offset timestamps. This enables precise control…

Sound · Computer Science 2026-02-10 Yisu Liu , Chenxing Li , Wanqian Zhang , Wenfu Wang , Meng Yu , Ruibo Fu , Zheng Lin , Weiping Wang , Dong Yu

With the rapid development of deep learning techniques, the generation and counterfeiting of multimedia material are becoming increasingly straightforward to perform. At the same time, sharing fake content on the web has become so simple…

Multimedia · Computer Science 2022-09-19 Davide Salvi , Brian Hosler , Paolo Bestagini , Matthew C. Stamm , Stefano Tubaro

As a fundamental task in long-form video understanding, temporal action detection (TAD) aims to capture inherent temporal relations in untrimmed videos and identify candidate actions with precise boundaries. Over the years, various…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Shuming Liu , Lin Sui , Chen-Lin Zhang , Fangzhou Mu , Chen Zhao , Bernard Ghanem

Recent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts. This improvement could be due to the powerful modeling ability of the multi-head…

Sound · Computer Science 2024-09-10 Duc-Tuan Truong , Ruijie Tao , Tuan Nguyen , Hieu-Thi Luong , Kong Aik Lee , Eng Siong Chng

This paper presents MAST, a new model for Multimodal Abstractive Text Summarization that utilizes information from all three modalities -- text, audio and video -- in a multimodal video. Prior work on multimodal abstractive text…

Computation and Language · Computer Science 2020-10-19 Aman Khullar , Udit Arora

Graph models provide efficient tools to capture the underlying structure of data defined over networks. Many real-world network topologies are subject to change over time. Learning to model the dynamic interactions between entities in such…

Machine Learning · Computer Science 2025-01-03 Amirhossein Javaheri , Jiaxi Ying , Daniel P. Palomar , Farokh Marvasti

Knowledge is inherently time-sensitive and continuously evolves over time. Although current Retrieval-Augmented Generation (RAG) systems enrich LLMs with external knowledge, they largely ignore this temporal nature. This raises two…

Information Retrieval · Computer Science 2025-10-16 Jiale Han , Austin Cheung , Yubai Wei , Zheng Yu , Xusheng Wang , Bing Zhu , Yi Yang

How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming learning that combining temporal and appearance information…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Jinxia Xie , Bineng Zhong , Qihua Liang , Ning Li , Zhiyi Mo , Shuxiang Song

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation,…

‹ Prev 1 3 4 5 6 7 10 Next ›