English
Related papers

Related papers: MonoSIM: An open source SIL framework for Ackerman…

200 papers

Sparse-LiDAR-prompted depth foundation models (PromptDA, Prior Depth Anything, DMD3C) have shown strong results on indoor scenes or within KITTI's standard 80-meter evaluation cap. However, two limitations remain: (i) systematic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Kai Zheng , Qiang Feng , Xingjian Liu , Wenquan Tan , Yuan Li

This paper presents a novel method to reduce the scale drift for indoor monocular simultaneous localization and mapping (SLAM). We leverage the prior knowledge that in the indoor environment, the line segments form tight clusters, e.g. many…

Computer Vision and Pattern Recognition · Computer Science 2018-11-06 Ting Sun , Dezhen Song , Dit-Yan Yeung , Ming Liu

Incremental Learning (IL) aims to learn new tasks while preserving previously acquired knowledge. Integrating the zero-shot learning capabilities of pre-trained vision-language models into IL methods has marked a significant advancement.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Haihua Luo , Xuming Ran , Jiangrong Shen , Timo Hämäläinen , Zhonghua Chen , Qi Xu , Fengyu Cong

End-to-end autonomous driving has advanced significantly, offering benefits such as system simplicity and stronger driving performance in both open-loop and closed-loop settings than conventional pipelines. However, existing frameworks…

Robotics · Computer Science 2025-06-04 Wei Liu , Jiyuan Zhang , Binxiong Zheng , Yufeng Hu , Yingzhan Lin , Zengfeng Zeng

Multimodal Large Language Models (MLLMs) achieve remarkable performance for fine-grained pixel-level understanding tasks. However, all the works rely heavily on extra components, such as vision encoder (CLIP), segmentation experts, leading…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Tao Zhang , Xiangtai Li , Zilong Huang , Yanwei Li , Weixian Lei , Xueqing Deng , Shihao Chen , Shunping Ji , Jiashi Feng

Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora, and rigorous evaluation. To address these challenges, we…

Computation and Language · Computer Science 2025-12-09 Bo Yang , Yunkui Chen , Lanfei Feng , Yu Zhang , Xiao Xu , Jianyu Zhang , Nueraili Aierken , Runhe Huang , Hongjian Lin , Yibin Ying , Shijian Li

The rapid growth of ride-sharing services presents a promising solution to urban transportation challenges, such as congestion and carbon emissions. However, developing efficient operational strategies, such as pricing, matching, and fleet…

Systems and Control · Electrical Eng. & Systems 2025-05-26 Wang Chen , Hongzheng Shi , Jintao Ke

A significant portion of driving hazards is caused by human error and disregard for local driving regulations; Consequently, an intelligent assistance system can be beneficial. This paper proposes a novel vision-based modular package to…

Computer Vision and Pattern Recognition · Computer Science 2023-03-30 Amirhossein Kazerouni , Amirhossein Heydarian , Milad Soltany , Aida Mohammadshahi , Abbas Omidi , Saeed Ebadollahi

In this paper, we address the novel, highly challenging problem of estimating the layout of a complex urban driving scenario. Given a single color image captured from a driving platform, we aim to predict the bird's-eye view layout of the…

Computer Vision and Pattern Recognition · Computer Science 2020-02-21 Kaustubh Mani , Swapnil Daga , Shubhika Garg , N. Sai Shankar , Krishna Murthy Jatavallabhula , K. Madhava Krishna

Traffic simulation is essential for autonomous vehicle (AV) development, enabling comprehensive safety evaluation across diverse driving conditions. However, traditional rule-based simulators struggle to capture complex human interactions,…

In this paper, we present an efficient visual SLAM system designed to tackle both short-term and long-term illumination challenges. Our system adopts a hybrid approach that combines deep learning techniques for feature detection and…

Robotics · Computer Science 2025-02-28 Kuan Xu , Yuefan Hao , Shenghai Yuan , Chen Wang , Lihua Xie

Automated driving technologies promise substantial improvements in transportation safety, efficiency, and accessibility. However, ensuring the reliability and safety of Autonomous Vehicles in complex, real-world environments remains a…

Software Engineering · Computer Science 2025-03-03 João-Vitor Zacchi , Edoardo Clementi , Núria Mata

This paper presents OpenREALM, a real-time mapping framework for Unmanned Aerial Vehicles (UAVs). A camera attached to the onboard computer of a moving UAV is utilized to acquire high resolution image mosaics of a targeted area of interest.…

Computer Vision and Pattern Recognition · Computer Science 2020-09-23 Alexander Kern , Markus Bobbe , Yogesh Khedar , Ulf Bestmann

Imitation learning has been a trend recently, yet training a generalist agent across multiple tasks still requires large-scale expert demonstrations, which are costly and labor-intensive to collect. To address the challenge of limited…

Robotics · Computer Science 2025-09-25 Yifan Ye , Jun Cen , Jing Chen , Zhihe Lu

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navigation simulators,…

Robotics · Computer Science 2024-10-15 Zihan Wang , Xiangyang Li , Jiahao Yang , Yeqi Liu , Shuqiang Jiang

We introduce Xmodel-VLM, a cutting-edge multimodal vision language model. It is designed for efficient deployment on consumer GPU servers. Our work directly confronts a pivotal industry issue by grappling with the prohibitive service costs…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Wanting Xu , Yang Liu , Langping He , Xucheng Huang , Ling Jiang

Increasing the implemented SAE level of autonomy in road vehicles requires extensive simulations and verifications in a realistic simulation environment before proving ground and public road testing. The level of detail in the simulation…

Robotics · Computer Science 2023-06-02 Mustafa Ridvan Cantas , Levent Guvenc

Our work introduces a module for assessing the trajectory safety of autonomous vehicles in dynamic environments marked by high uncertainty. We focus on occluded areas and occluded traffic participants with limited information about…

Robotics · Computer Science 2024-07-31 Korbinian Moller , Rainer Trauth , Johannes Betz

Multimodal large language models (MLLMs) achieve strong performance by jointly processing inputs from multiple modalities, such as vision, audio, and language. However, building such models or extending them to new modalities often requires…

Machine Learning · Computer Science 2026-03-24 Md Kaykobad Reza , Ameya Patil , Edward Ayrapetian , M. Salman Asif

The lane-level localization accuracy is very important for autonomous vehicles. The Global Navigation Satellite System (GNSS), e.g. GPS, is a generic localization method for vehicles, but is vulnerable to the multi-path interference in the…

Computer Vision and Pattern Recognition · Computer Science 2018-08-24 Iljoo Baek , Mengwen He
‹ Prev 1 8 9 10 Next ›