English
Related papers

Related papers: EIMC: Efficient Instance-aware Multi-modal Collabo…

200 papers

Real-world Vehicle-to-Everything (V2X) cooperative perception systems often operate under heterogeneous sensor configurations due to cost constraints and deployment variability across vehicles and infrastructure. This heterogeneity poses…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Chuheng Wei , Ziye Qin , Walter Zimmer , Guoyuan Wu , Matthew J. Barth

Cross-modal transformers have demonstrated superiority in various vision tasks by effectively integrating different modalities. This paper first critiques prior token exchange methods which replace less informative tokens with inter-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Ding Jia , Jianyuan Guo , Kai Han , Han Wu , Chao Zhang , Chang Xu , Xinghao Chen

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited number of works have…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 Alejandro Cartas , Jordi Luque , Petia Radeva , Carlos Segura , Mariella Dimiccoli

Cooperative perception extends the perception capabilities of autonomous vehicles by enabling multi-agent information sharing via Vehicle-to-Everything (V2X) communication. Unlike traditional onboard sensors, V2X acts as a dynamic…

Other Computer Science · Computer Science 2025-05-05 Zhiying Song , Tenghui Xie , Fuxi Wen , Jun Li

Vehicular communication systems operating in the millimeter wave (mmWave) band are highly susceptible to signal blockage from dynamic obstacles such as vehicles, pedestrians, and infrastructure. To address this challenge, we propose a…

Machine Learning · Computer Science 2025-07-22 Ahmad M. Nazar , Abdulkadir Celik , Mohamed Y. Selim , Asmaa Abdallah , Daji Qiao , Ahmed M. Eltawil

Human drivers adeptly navigate complex scenarios by utilizing rich attentional semantics, but the current autonomous systems struggle to replicate this ability, as they often lose critical semantic information when converting 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Pei Liu , Haipeng Liu , Haichao Liu , Xin Liu , Jinxin Ni , Jun Ma

Multimodal fingerprinting is a crucial technique to sub-meter 6G integrated sensing and communications (ISAC) localization, but two hurdles block deployment: (i) the contribution each modality makes to the target position varies with the…

The integration of multimodal sensing and millimeter-wave (mmWave) communications is a key enabler for highly mobile vehicle-to-infrastructure (V2I) networks. However, continuous high-resolution visual sensing incurs prohibitive…

Signal Processing · Electrical Eng. & Systems 2026-04-09 Wenqi Fan , Ning Wei , Rongyan Xi , Ahmad Bazzi , Yue Xiu , Chadi Assi , Jing Dong , Jing Jin

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Yikai Wang , Xinghao Chen , Lele Cao , Wenbing Huang , Fuchun Sun , Yunhe Wang

Vehicle-Infrastructure Collaborative Perception (VICP) is pivotal for resolving occlusion in autonomous driving, yet the trade-off between communication bandwidth and feature redundancy remains a critical bottleneck. While intermediate…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Li Wang , Boqi Li , Hang Chen , Xingjian Wu , Yichen Wang , Jiewen Tan , Xinyu Zhang , Huaping Liu

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha

Multi-agent collaborative perception could significantly upgrade the perception performance by enabling agents to share complementary information with each other through communication. It inevitably results in a fundamental trade-off…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Yue Hu , Shaoheng Fang , Zixing Lei , Yiqi Zhong , Siheng Chen

Multimodal speech emotion recognition aims to detect speakers' emotions from audio and text. Prior works mainly focus on exploiting advanced networks to model and fuse different modality information to facilitate performance, while…

Computation and Language · Computer Science 2023-04-11 Zhen Wu , Yizhe Lu , Xinyu Dai

The Mixture-of-Experts (MoE) paradigm has emerged as a promising solution to scale up model capacity while maintaining inference efficiency. However, deploying MoE models across heterogeneous end-cloud environments poses new challenges in…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-08-11 Zheming Yang , Yunqing Hu , Sheng Sun , Wen Ji

Beamforming techniques are utilized in millimeter wave (mmWave) communication to address the inherent path loss limitation, thereby establishing and maintaining reliable connections. However, adopting standard defined beamforming approach…

Networking and Internet Architecture · Computer Science 2025-09-16 Muhammad Baqer Mollah , Honggang Wang , Hua Fang

Cooperative perception aims to address the inherent limitations of single-vehicle autonomous driving systems through information exchange among multiple agents. Previous research has primarily focused on single-frame perception tasks.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Jiaru Zhong , Jiahao Wang , Jiahui Xu , Xiaofan Li , Zaiqing Nie , Haibao Yu

Multi-modal fusion is a basic task of autonomous driving system perception, which has attracted many scholars' interest in recent years. The current multi-modal fusion methods mainly focus on camera data and LiDAR data, but pay little…

Robotics · Computer Science 2022-11-14 Yan Gong , Jianli Lu , Jiayi Wu , Wenzhuo Liu

The objective of the collaborative vehicle-to-everything perception task is to enhance the individual vehicle's perception capability through message communication among neighboring traffic agents. Previous methods focus on achieving…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Si Liu , Zihan Ding , Jiahui Fu , Hongyu Li , Siheng Chen , Shifeng Zhang , Xu Zhou

Learning effective joint embedding for cross-modal data has always been a focus in the field of multimodal machine learning. We argue that during multimodal fusion, the generated multimodal embedding may be redundant, and the discriminative…

Machine Learning · Computer Science 2022-12-06 Sijie Mai , Ying Zeng , Haifeng Hu

The research and applications of multimodal emotion recognition have become increasingly popular recently. However, multimodal emotion recognition faces the challenge of lack of data. To solve this problem, we propose to use transfer…

Computation and Language · Computer Science 2022-07-13 Zihan Zhao , Yanfeng Wang , Yu Wang
‹ Prev 1 3 4 5 6 7 10 Next ›