English
Related papers

Related papers: LFTag: A Scalable Visual Fiducial System with Low …

200 papers

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field…

Computer Vision and Pattern Recognition · Computer Science 2022-05-19 Pengyuan Wang , HyunJun Jung , Yitong Li , Siyuan Shen , Rahul Parthasarathy Srikanth , Lorenzo Garattoni , Sven Meier , Nassir Navab , Benjamin Busam

Vision Transformers have achieved impressive performance in video classification, while suffering from the quadratic complexity caused by the Softmax attention mechanism. Some studies alleviate the computational costs by reducing the number…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Kaiyue Lu , Zexiang Liu , Jianyuan Wang , Weixuan Sun , Zhen Qin , Dong Li , Xuyang Shen , Hui Deng , Xiaodong Han , Yuchao Dai , Yiran Zhong

Concurrently estimating the 6-DOF pose of multiple cameras or robots---cooperative localization---is a core problem in contemporary robotics. Current works focus on a set of mutually observable world landmarks and often require inbuilt…

Robotics · Computer Science 2013-04-03 Vikas Dhiman , Julian Ryde , Jason J. Corso

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data,…

Robotics · Computer Science 2025-10-20 Fuhao Li , Wenxuan Song , Han Zhao , Jingbo Wang , Pengxiang Ding , Donglin Wang , Long Zeng , Haoang Li

Scalable and maintainable map representations are fundamental to enabling large-scale visual navigation and facilitating the deployment of robots in real-world environments. While collaborative localization across multi-session mapping…

Robotics · Computer Science 2026-01-21 Jianhao Jiao , Changkun Liu , Jingwen Yu , Boyi Liu , Qianyi Zhang , Yue Wang , Dimitrios Kanoulas

Document Visual Question Answering (DocVQA) requires models to jointly understand textual semantics, spatial layout, and visual features. Current methods struggle with explicit spatial relationship modeling, inefficiency with…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ahmad Mohammadshirazi , Pinaki Prasad Guha Neogi , Dheeraj Kulshrestha , Rajiv Ramnath

Small object detection remains a significant challenge due to feature degradation from downsampling, mutual occlusion in dense clusters, and complex background interference. To address these issues, this paper proposes FSDETR, a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Jianchao Huang , Fengming Zhang , Haibo Zhu , Tao Yan

We propose a new topological tool for computer vision - Scalar Function Topology Divergence (SFTD), which measures the dissimilarity of multi-scale topology between sublevel sets of two functions having a common domain. Functions can be…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Ilya Trofimov , Daria Voronkova , Eduard Tulchinskii , Evgeny Burnaev , Serguei Barannikov

With the rapid development of large multimodal models (LMMs), multimodal understanding applications are emerging. As most LMM inference requests originate from edge devices with limited computational capabilities, the predominant inference…

Signal Processing · Electrical Eng. & Systems 2025-11-05 Cheng Yuan , Zhening Liu , Jiashu Lv , Jiawei Shao , Yufei Jiang , Jun Zhang , Xuelong Li

Indoor localization of autonomous mobile robots (AMRs) can be realized with fiducial markers. Such systems require only a simple, monocular camera as sensor and fiducial markers as passive, identifiable position references that can be…

Image and Video Processing · Electrical Eng. & Systems 2026-01-27 Sven Hinderer , Martina Scheffler , Bin Yang

Fiducial markers are a computer vision tool used for object pose estimation and detection. These markers are highly useful in fields such as industry, medicine and logistics. However, optimal lighting conditions are not always available,and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Rafael Berral-Soler , Rafael Muñoz-Salinas , Rafael Medina-Carnicer , Manuel J. Marín-Jiménez

Light detection and ranging (LiDAR) systems are pivotal for precise distance and velocity measurement, yet widespread deployment requires solutions that balance their performance, robustness, and simplicity. Here, we propose a novel chaotic…

Optics · Physics 2026-02-17 T. Wang , Z. Li , H. Shen , Y. Ma , Y. Li , S. Xiang , S. Baland , Y. Hao

This paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S$^3$FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we…

Computer Vision and Pattern Recognition · Computer Science 2017-11-16 Shifeng Zhang , Xiangyu Zhu , Zhen Lei , Hailin Shi , Xiaobo Wang , Stan Z. Li

Globally consistent dense maps are a key requirement for long-term robot navigation in complex environments. While previous works have addressed the challenges of dense mapping and global consistency, most require more computational…

Visual localization is a fundamental task for various applications including autonomous driving and robotics. Prior methods focus on extracting large amounts of often redundant locally reliable features, resulting in limited efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Fei Xue , Ignas Budvytis , Roberto Cipolla

We propose a fixed-lag smoother-based sensor fusion architecture to leverage the complementary benefits of range-based sensors and visual-inertial odometry (VIO) for localization. We use two fixed-lag smoothers (FLS) to decouple accurate…

Robotics · Computer Science 2024-01-05 Abhishek Goudar , Wenda Zhao , Angela P. Schoellig

Precise segmentation of retinal arteries and veins carries the diagnosis of systemic cardiovascular conditions. However, standard convolutional architectures often yield topologically disjointed segmentations, characterized by gaps and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Iftekhar Ahmed , Shakib Absar , Aftar Ahmad Sami , Shadman Sakib , Debojyoti Biswas , Seraj Al Mahmud Mostafa

We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-scale dataset, Mono3DRefer, which contains 3D object targets…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Yang Zhan , Yuan Yuan , Zhitong Xiong

3D lane detection and topology reasoning are essential tasks in autonomous driving scenarios, requiring not only detecting the accurate 3D coordinates on lane lines, but also reasoning the relationship between lanes and traffic elements.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Han Li , Zehao Huang , Zitian Wang , Wenge Rong , Naiyan Wang , Si Liu

Vision-language pre-training (VLP) has emerged as an effective scheme for multimodal representation learning, but its reliance on large-scale multimodal data poses significant challenges for medical applications. Federated learning (FL)…

Machine Learning · Computer Science 2024-11-25 Zitao Shuai , Chenwei Wu , Zhengxu Tang , Liyue Shen