English
Related papers

Related papers: Pilot-guided Multimodal Semantic Communication for…

200 papers

In this paper, we introduce a new problem, Online-MMSI, where the model must perform multimodal social interaction understanding (MMSI) using only historical information. Given a recorded video and a multi-party dialogue, the AI assistant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xinpeng Li , Shijian Deng , Bolin Lai , Weiguo Pian , James M. Rehg , Yapeng Tian

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Ivan Kukanov , Jun Wah Ng

The problem of goal-oriented semantic filtering and timely source coding in multiuser communication systems is considered here. We study a distributed monitoring system in which multiple information sources, each observing a physical…

Information Theory · Computer Science 2024-02-16 Pouya Agheli , Nikolaos Pappas , Marios Kountouris

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-24 Ya Jiang , Qing Wang , Jun Du , Maocheng Hu , Pengfei Hu , Zeyan Liu , Shi Cheng , Zhaoxu Nian , Yuxuan Dong , Mingqi Cai , Xin Fang , Chin-Hui Lee

6G system is evolving toward full-spectrum coverage,ultra-wide bandwidth, and high mobility, resulting in increasingly complex propagation environments. The deep integration of communication and sensing is widely recognized as a core 6G…

Information Theory · Computer Science 2026-01-27 Xuejian Zhang , Ruisi He , Mi Yang , Zhengyu Zhang , Ziyi Qi

Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as conditioning inputs to generate motions that align with the…

Sound · Computer Science 2025-05-30 Zi-An Wang , Shihao Zou , Shiyao Yu , Mingyuan Zhang , Chao Dong

Accurate perception of dynamic traffic scenes is crucial for high-level autonomous driving systems, requiring robust object motion estimation and instance segmentation. However, traditional methods often treat them as separate tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Yinqi Chen , Meiying Zhang , Qi Hao , Guang Zhou

Videos are a rich source of multi-modal supervision. In this work, we learn representations using self-supervision by leveraging three modalities naturally present in videos: visual, audio and language streams. To this end, we introduce the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-02 Jean-Baptiste Alayrac , Adrià Recasens , Rosalia Schneider , Relja Arandjelović , Jason Ramapuram , Jeffrey De Fauw , Lucas Smaira , Sander Dieleman , Andrew Zisserman

Molecular communication (MC) provides a foundational framework for information transmission in the Internet of Bio-Nano Things (IoBNT), where efficiency and reliability are crucial. However, the inherent limitations of molecular channels,…

Signal Processing · Electrical Eng. & Systems 2025-04-02 Hanlin Cai , Ozgur B. Akan

This paper investigates robust semantic communications over multiple-input multiple-output (MIMO) fading channels. Current semantic communications over MIMO channels mainly focus on channel adaptive encoding and decoding, which lacks…

Information Theory · Computer Science 2024-07-09 Yiheng Duan , Tong Wu , Zhiyong Chen , Meixia Tao

Advancements in prompt tuning of vision-language models have underscored their potential in enhancing open-world visual concept comprehension. However, prior works only primarily focus on single-mode (only one prompt for each modality) and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Dongsheng Wang , Miaoge Li , Xinyang Liu , MingSheng Xu , Bo Chen , Hanwang Zhang

We consider a multi-user semantic communications system in which agents (transmitters and receivers) interact through the exchange of semantic messages to convey meanings. In this context, languages are instrumental in structuring the…

Artificial Intelligence · Computer Science 2023-08-09 Mohamed Sana , Emilio Calvanese Strinati

The event-based Vision-Language Model (VLM) recently has made good progress for practical vision tasks. However, most of these works just utilize CLIP for focusing on traditional perception tasks, which obstruct model understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Pengteng Li , Yunfan Lu , Pinghao Song , Wuyang Li , Huizai Yao , Hui Xiong

Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. Current multi-modal approaches mainly focus on integrating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jinkun Dai , Yuanxin Ye , Peng Tang , Tengfeng Tang , Xianping Ma , Jing Xiao , Mi Wang

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Current optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel in challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Qianang Zhou , Junhui Hou , Meiyi Yang , Yongjian Deng , Youfu Li , Junlin Xiong

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Vladimir Iashin , Esa Rahtu

This paper presents Multimodal-Wireless, a large-scale open-source dataset for multimodal sensing and communication research. The dataset is generated through an integrated and customizable data pipeline built upon the CARLA simulator and…

Signal Processing · Electrical Eng. & Systems 2026-02-12 Tianhao Mao , Le Liang , Jie Yang , Hao Ye , Shi Jin , Geoffrey Ye Li

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic

Multimodal semantic communication has gained widespread attention due to its ability to enhance downstream task performance. A key challenge in such systems is the effective fusion of features from different modalities, which requires the…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Haoshuo Zhang , Yufei Bo , Hongwei Zhang , Meixia Tao
‹ Prev 1 4 5 6 7 8 10 Next ›