English
Related papers

Related papers: CoCMT: Communication-Efficient Cross-Modal Transfo…

200 papers

Collaborative Perception (CP) has been a promising solution to address occlusions in the traffic environment by sharing sensor data among collaborative vehicles (CoV) via vehicle-to-everything (V2X) network. With limited wireless bandwidth,…

Robotics · Computer Science 2024-07-02 Yukuan Jia , Yuxuan Sun , Ruiqing Mao , Zhaojun Nan , Sheng Zhou , Zhisheng Niu

Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerful graph transformers to efficiently model the spatial and…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Peng Chu , Jiang Wang , Quanzeng You , Haibin Ling , Zicheng Liu

The transformer model has gained widespread adoption in computer vision tasks in recent times. However, due to the quadratic time and memory complexity of self-attention, which is proportional to the number of input tokens, most existing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Wei Tan , Yifeng Geng , Xuansong Xie

Today, the acquisition of various behavioral log data has enabled deeper understanding of customer preferences and future behaviors in the marketing field. In particular, multimodal deep learning has achieved highly accurate predictions by…

Computational Engineering, Finance, and Science · Computer Science 2024-05-14 Junichiro Niimi

Hybrid vision architectures combining Transformers and CNNs have significantly advanced image classification, but they usually do so at significant computational cost. We introduce EVCC (Enhanced Vision Transformer-ConvNeXt-CoAtNet), a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Kazi Reyazul Hasan , Md Nafiu Rahman , Wasif Jalal , Sadif Ahmed , Shahriar Raj , Mubasshira Musarrat , Muhammad Abdullah Adnan

Multimodal emotion recognition (MER) aims to infer human affect by jointly modeling audio and visual cues; however, existing approaches often struggle with temporal misalignment, weakly discriminative feature representations, and suboptimal…

Multimedia · Computer Science 2026-01-21 Joe Dhanith P R , Shravan Venkatraman , Vigya Sharma , Santhosh Malarvannan

Collaborative perception has garnered considerable attention due to its capacity to address several inherent challenges in single-agent perception, including occlusion and out-of-range issues. However, existing collaborative perception…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Zhenyang Ni , Zixing Lei , Yifan Lu , Dingju Wang , Chen Feng , Yanfeng Wang , Siheng Chen

Transformer is a powerful model for text understanding. However, it is inefficient due to its quadratic complexity to input sequence length. Although there are many methods on Transformer acceleration, they are still either inefficient on…

Computation and Language · Computer Science 2021-09-07 Chuhan Wu , Fangzhao Wu , Tao Qi , Yongfeng Huang , Xing Xie

Cooperative perception significantly enhances scene understanding by integrating complementary information from diverse agents. However, existing research often overlooks critical challenges inherent in real-world multi-source data…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Gong Chen , Chaokun Zhang , Tao Tang , Pengcheng Lv , Feng Li , Xin Xie

Situational awareness as a necessity in the connected and autonomous vehicles (CAV) domain is the subject of a significant number of researches in recent years. The driver's safety is directly dependent on the robustness, reliability, and…

Computer Vision and Pattern Recognition · Computer Science 2020-10-23 Ehsan Emad Marvasti , Arash Raftari , Amir Emad Marvasti , Yaser P. Fallah

Image-level weakly supervised semantic segmentation is a challenging task that has been deeply studied in recent years. Most of the common solutions exploit class activation map (CAM) to locate object regions. However, such response maps…

Computer Vision and Pattern Recognition · Computer Science 2023-10-02 Yukun Su , Jingliang Deng , Zonghan Li

In recent years, autonomous driving has garnered significant attention due to its potential for improving road safety through collaborative perception among connected and autonomous vehicles (CAVs). However, time-varying channel variations…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Yuang Zhang , Haonan An , Zhengru Fang , Guowen Xu , Yuan Zhou , Xianhao Chen , Yuguang Fang

Existing Vehicle-to-Everything (V2X) cooperative perception methods rely on accurate multi-agent 3D annotations. Nevertheless, it is time-consuming and expensive to collect and annotate real-world data, especially for V2X systems. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Seth Z. Zhao , Hao Xiang , Chenfeng Xu , Xin Xia , Bolei Zhou , Jiaqi Ma

Connected and Automated Vehicles (CAVs) utilize a variety of onboard sensors to sense their surrounding environment. CAVs can improve their perception capabilities if vehicles exchange information about what they sense using V2X…

Networking and Internet Architecture · Computer Science 2020-11-11 Gokulnath Thandavarayan , Miguel Sepulcre , Javier Gozalvez

Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Shivansh Patel , Saim Wani , Unnat Jain , Alexander Schwing , Svetlana Lazebnik , Manolis Savva , Angel X. Chang

Perceiving the complex driving environment precisely is crucial to the safe operation of autonomous vehicles. With the tremendous advancement of deep learning and communication technology, Vehicle-to-Everything (V2X) collaboration has the…

Software Engineering · Computer Science 2024-08-30 An Guo , Xinyu Gao , Zhenyu Chen , Yuan Xiao , Jiakai Liu , Xiuting Ge , Weisong Sun , Chunrong Fang

Transparent objects, such as glass walls and doors, constitute architectural obstacles hindering the mobility of people with low vision or blindness. For instance, the open space behind glass doors is inaccessible, unless it is correctly…

Computer Vision and Pattern Recognition · Computer Science 2021-08-23 Jiaming Zhang , Kailun Yang , Angela Constantinescu , Kunyu Peng , Karin Müller , Rainer Stiefelhagen

Autonomous driving relies on accurate perception to ensure safe driving. Collaborative perception improves accuracy by mitigating the sensing limitations of individual vehicles, such as limited perception range and occlusion-induced blind…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-21 Hui Zhang , Yuquan Yang , Zechuan Gong , Xiaohua Xu , Dan Keun Sung

Audio-guided Video Object Segmentation (A-VOS) and Referring Video Object Segmentation (R-VOS) are two highly related tasks that both aim to segment specific objects from video sequences according to expression prompts. However, due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Jiajun Chen , Jiacheng Lin , Guojin Zhong , Haolong Fu , Ke Nai , Kailun Yang , Zhiyong Li

Extracting robust feature representation is critical for object re-identification to accurately identify objects across non-overlapping cameras. Although having a strong representation ability, the Vision Transformer (ViT) tends to overfit…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Lei Tan , Pingyang Dai , Jie Chen , Liujuan Cao , Yongjian Wu , Rongrong Ji