中文
相关论文

相关论文: PaveCap: The First Multimodal Framework for Compre…

200 篇论文

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu

Fully Convolution Networks (FCN) have achieved great success in dense prediction tasks including semantic segmentation. In this paper, we start from discussing FCN by understanding its architecture limitations in building a strong…

计算机视觉与模式识别 · 计算机科学 2016-11-29 Bing Shuai , Ting Liu , Gang Wang

3D dense captioning requires a model to translate its understanding of an input 3D scene into several captions associated with different object regions. Existing methods adopt a sophisticated "detect-then-describe" pipeline, which builds…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Sijin Chen , Hongyuan Zhu , Mingsheng Li , Xin Chen , Peng Guo , Yinjie Lei , Gang Yu , Taihao Li , Tao Chen

Real-time pavement condition monitoring provides highway agencies with timely and accurate information that could form the basis of pavement maintenance and rehabilitation policies. Existing technologies rely heavily on manual data…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Abdulateef Daud , Mark Amo-Boateng , Neema Jakisa Owor , Armstrong Aboah , Yaw Adu-Gyamfi

In this paper, we propose a learning-based method for predicting dense depth values of a scene from a monocular omnidirectional image. An omnidirectional image has a full field-of-view, providing much more complete descriptions of the scene…

计算机视觉与模式识别 · 计算机科学 2022-02-09 Jiayang Bai , Shuichang Lai , Haoyu Qin , Jie Guo , Yanwen Guo

3D dense captioning aims to describe individual objects by natural language in 3D scenes, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous…

计算机视觉与模式识别 · 计算机科学 2022-04-07 Zhihao Yuan , Xu Yan , Yinghong Liao , Yao Guo , Guanbin Li , Zhen Li , Shuguang Cui

Achieving reliable multidimensional Vehicle-to-Vehicle (V2V) channel state information (CSI) prediction is both challenging and crucial for optimizing downstream tasks that depend on instantaneous CSI. This work extends traditional…

系统与控制 · 电气工程与系统科学 2024-09-24 Lei Chu , Daoud Burghal , Rui Wang , Michael Neuman , Andreas F. Molisch

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Yunxiang Yang , Ningning Xu , Jidong J. Yang

Understanding the traversability of terrain is essential for autonomous robot navigation, particularly in unstructured environments such as natural landscapes. Although traditional methods, such as occupancy mapping, provide a basic…

While most prior research has focused on improving the precision of multimodal trajectory predictions, the explicit modeling of multimodal behavioral intentions (e.g., yielding, overtaking) remains relatively underexplored. This paper…

机器人学 · 计算机科学 2026-02-19 Jiawei Sun , Xibin Yue , Jiahui Li , Tianle Shen , Chengran Yuan , Shuo Sun , Sheng Guo , Quanyun Zhou , Marcelo H Ang

Predicting future trajectories of road agents is a critical task for autonomous driving. Recent goal-based trajectory prediction methods, such as DenseTNT and PECNet, have shown good performance on prediction tasks on public datasets.…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Qiujing Lu , Weiqiao Han , Jeffrey Ling , Minfa Wang , Haoyu Chen , Balakrishnan Varadarajan , Paul Covington

Road transport infrastructure is critical for safe, fast, economical, and reliable mobility within the whole country that is conducive to a productive society. However, roads tend to deteriorate over time due to natural causes in the…

计算机视觉与模式识别 · 计算机科学 2021-03-12 James-Andrew Sarmiento

We propose Dense FixMatch, a simple method for online semi-supervised learning of dense and structured prediction tasks combining pseudo-labeling and consistency regularization via strong data augmentation. We enable the application of…

计算机视觉与模式识别 · 计算机科学 2022-10-19 Miquel Martí i Rabadán , Alessandro Pieropan , Hossein Azizpour , Atsuto Maki

Explainability and transparent decision-making are essential for the safe deployment of autonomous driving systems. Scene captioning summarizes environmental conditions and risk factors in natural language, improving transparency, safety,…

机器人学 · 计算机科学 2026-03-03 Zihang Wang , Xu Li , Benwu Wang , Wenkai Zhu , Xieyuanli Chen , Dong Kong , Kailin Lyu , Yinan Du , Yiming Peng , Haoyang Che

Having good knowledge of terrain information is essential for improving the performance of various downstream tasks on complex terrains, especially for the locomotion and navigation of legged robots. We present a novel framework for neural…

机器人学 · 计算机科学 2024-03-13 Bowen Yang , Qingwen Zhang , Ruoyu Geng , Lujia Wang , Ming Liu

When designing a semantic segmentation module for a practical application, such as autonomous driving, it is crucial to understand the robustness of the module with respect to a wide range of image corruptions. While there are recent…

计算机视觉与模式识别 · 计算机科学 2020-08-26 Christoph Kamann , Carsten Rother

Weather is an important factor affecting transportation and road safety. In this paper, we leverage state-of-the-art convolutional neural networks in labelling images taken by street and highway cameras located across across North America.…

计算机视觉与模式识别 · 计算机科学 2020-01-28 Sheela Ramanna , Cenker Sengoz , Scott Kehler , Dat Pham

Understanding the complex urban infrastructure with centimeter-level accuracy is essential for many applications from autonomous driving to mapping, infrastructure monitoring, and urban management. Aerial images provide valuable information…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Seyed Majid Azimi , Corentin Henry , Lars Sommer , Arne Schumann , Eleonora Vig

Accurate perception of dynamic traffic scenes is crucial for high-level autonomous driving systems, requiring robust object motion estimation and instance segmentation. However, traditional methods often treat them as separate tasks,…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Yinqi Chen , Meiying Zhang , Qi Hao , Guang Zhou

CLIP models perform remarkably well on zero-shot classification and retrieval tasks. But recent studies have shown that learnt representations in CLIP are not well suited for dense prediction tasks like object detection, semantic…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Pavan Kumar Anasosalu Vasu , Hadi Pouransari , Fartash Faghri , Oncel Tuzel