中文
相关论文

相关论文: CoCMT: Communication-Efficient Cross-Modal Transfo…

200 篇论文

Semantic segmentation assigns labels to pixels in images, a critical yet challenging task in computer vision. Convolutional methods, although capturing local dependencies well, struggle with long-range relationships. Vision Transformers…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Mian Muhammad Naeem Abid , Nancy Mehta , Zongwei Wu , Radu Timofte

Collaborative perception is vital for autonomous driving yet remains constrained by tight communication budgets. Earlier work reduced bandwidth by compressing full feature maps with fixed-rate encoders, which adapts poorly to a changing…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Gong Chen , Chaokun Zhang , Xinyan Zhao

We propose a novel cascaded cross-modal transformer (CCMT) that combines speech and text transcripts to detect customer requests and complaints in phone conversations. Our approach leverages a multimodal paradigm by transcribing the speech…

计算与语言 · 计算机科学 2023-07-31 Nicolae-Catalin Ristea , Radu Tudor Ionescu

In this paper, we propose the problem of collaborative perception, where robots can combine their local observations with those of neighboring agents in a learnable way to improve accuracy on a perception task. Unlike existing work in…

计算机视觉与模式识别 · 计算机科学 2020-03-24 Yen-Cheng Liu , Junjiao Tian , Chih-Yao Ma , Nathan Glaser , Chia-Wen Kuo , Zsolt Kira

Autonomous Vehicles (AVs) rely on individual perception systems to navigate safely. However, these systems face significant challenges in adverse weather conditions, complex road geometries, and dense traffic scenarios. Cooperative…

机器人学 · 计算机科学 2025-03-25 Ahmad Sarlak , Rahul Amin , Abolfazl Razi

Vehicle-to-Infrastructure (V2I) collaborative perception leverages data collected by infrastructure's sensors to enhance vehicle perceptual capabilities. LiDAR, as a commonly used sensor in cooperative perception, is widely equipped in…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Xinxin Feng , Haoran Sun , Haifeng Zheng

Collaborative perception, an emerging paradigm in autonomous driving, has been introduced to mitigate the limitations of single-vehicle systems, such as limited sensor range and occlusion. To improve the robustness of inter-vehicle data…

信号处理 · 电气工程与系统科学 2025-11-26 Mingyi Lu , Guowei Liu , Le Liang , Chongtao Guo , Hao Ye , Shi Jin

Following their success in natural language processing, transformers have recently shown much promise for computer vision. The self-attention operation underlying transformers yields global interactions between all tokens ,i.e. words or…

Achieving fully autonomous driving with enhanced safety and efficiency relies on vehicle-to-everything cooperative perception, which enables vehicles to share perception data, thereby enhancing situational awareness and overcoming the…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Tao Huang , Jianan Liu , Xi Zhou , Dinh C. Nguyen , Mostafa Rahimi Azghadi , Yuxuan Xia , Qing-Long Han , Sumei Sun

The confluence of the advancement of Autonomous Vehicles (AVs) and the maturity of Vehicle-to-Everything (V2X) communication has enabled the capability of cooperative connected and automated vehicles (CAVs). Building on top of cooperative…

机器人学 · 计算机科学 2025-03-14 Zehao Wang , Yuping Wang , Zhuoyuan Wu , Hengbo Ma , Zhaowei Li , Hang Qiu , Jiachen Li

The embodied intelligence bridges the physical world and information space. As its typical physical embodiment, humanoid robots have shown great promise through robot learning algorithms in recent years. In this study, a hardware platform,…

机器人学 · 计算机科学 2025-10-17 Jiaxin Huang , Hanyu Liu , Yunsheng Ma , Jian Shen , Yilin Zheng , Jiayi Wen , Baishu Wan , Pan Li , Zhigong Song

This paper addresses the task of joint multi-agent perception and planning, especially as it relates to the real-world challenge of collision-free navigation for connected self-driving vehicles. For this task, several communication-enabled…

机器人学 · 计算机科学 2023-03-13 Nathaniel Moore Glaser , Zsolt Kira

Camera-only 3D detection provides an economical solution with a simple configuration for localizing objects in 3D space compared to LiDAR-based detection systems. However, a major challenge lies in precise depth estimation due to the lack…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Yue Hu , Yifan Lu , Runsheng Xu , Weidi Xie , Siheng Chen , Yanfeng Wang

This paper presents Camera-LiDAR Fusion Transformer (CLFT) models for traffic object segmentation, which leverage the fusion of camera and LiDAR data using vision transformers. Building on the methodology of visual transformers that exploit…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Toomas Tahves , Junyi Gu , Mauro Bellone , Raivo Sell

The recent trend in multiple object tracking (MOT) is heading towards leveraging deep learning to boost the tracking performance. In this paper, we propose a novel solution named TransSTAM, which leverages Transformer to effectively model…

计算机视觉与模式识别 · 计算机科学 2022-06-01 Peng Dai , Yiqiang Feng , Renliang Weng , Changshui Zhang

Vehicle-to-infrastructure (V2I) cooperative perception plays a crucial role in autonomous driving scenarios. Despite its potential to improve perception accuracy and robustness, the large amount of raw sensor data inevitably results in high…

信号处理 · 电气工程与系统科学 2024-07-31 Jiawei Shao , Teng Li , Jun Zhang

A fine-grained understanding of egocentric human-environment interactions is crucial for developing next-generation embodied agents. One fundamental challenge in this area involves accurately parsing hands and active objects. While…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Yuejiao Su , Yi Wang , Lei Yao , Yawen Cui , Lap-Pui Chau

Vision transformers have been successfully applied to image recognition tasks due to their ability to capture long-range dependencies within an image. However, there are still gaps in both performance and computational cost between…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Jianyuan Guo , Kai Han , Han Wu , Yehui Tang , Xinghao Chen , Yunhe Wang , Chang Xu

Collaborative perception enables more accurate and comprehensive scene understanding by learning how to share information between agents, with LiDAR point clouds providing essential precise spatial data. Due to the substantial data volume…

信号处理 · 电气工程与系统科学 2025-09-09 Ensong Liu , Rongqing Zhang , Xiang Cheng , Jian Tang

Entropy modeling is a key component for high-performance image compression algorithms. Recent developments in autoregressive context modeling helped learning-based methods to surpass their classical counterparts. However, the performance of…

图像与视频处理 · 电气工程与系统科学 2024-02-28 A. Burakhan Koyuncu , Han Gao , Atanas Boev , Georgii Gaikov , Elena Alshina , Eckehard Steinbach