中文
相关论文

相关论文: LFTag: A Scalable Visual Fiducial System with Low …

200 篇论文

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field…

计算机视觉与模式识别 · 计算机科学 2022-05-19 Pengyuan Wang , HyunJun Jung , Yitong Li , Siyuan Shen , Rahul Parthasarathy Srikanth , Lorenzo Garattoni , Sven Meier , Nassir Navab , Benjamin Busam

Vision Transformers have achieved impressive performance in video classification, while suffering from the quadratic complexity caused by the Softmax attention mechanism. Some studies alleviate the computational costs by reducing the number…

计算机视觉与模式识别 · 计算机科学 2022-10-18 Kaiyue Lu , Zexiang Liu , Jianyuan Wang , Weixuan Sun , Zhen Qin , Dong Li , Xuyang Shen , Hui Deng , Xiaodong Han , Yuchao Dai , Yiran Zhong

Concurrently estimating the 6-DOF pose of multiple cameras or robots---cooperative localization---is a core problem in contemporary robotics. Current works focus on a set of mutually observable world landmarks and often require inbuilt…

机器人学 · 计算机科学 2013-04-03 Vikas Dhiman , Julian Ryde , Jason J. Corso

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data,…

机器人学 · 计算机科学 2025-10-20 Fuhao Li , Wenxuan Song , Han Zhao , Jingbo Wang , Pengxiang Ding , Donglin Wang , Long Zeng , Haoang Li

Scalable and maintainable map representations are fundamental to enabling large-scale visual navigation and facilitating the deployment of robots in real-world environments. While collaborative localization across multi-session mapping…

机器人学 · 计算机科学 2026-01-21 Jianhao Jiao , Changkun Liu , Jingwen Yu , Boyi Liu , Qianyi Zhang , Yue Wang , Dimitrios Kanoulas

Document Visual Question Answering (DocVQA) requires models to jointly understand textual semantics, spatial layout, and visual features. Current methods struggle with explicit spatial relationship modeling, inefficiency with…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Ahmad Mohammadshirazi , Pinaki Prasad Guha Neogi , Dheeraj Kulshrestha , Rajiv Ramnath

Small object detection remains a significant challenge due to feature degradation from downsampling, mutual occlusion in dense clusters, and complex background interference. To address these issues, this paper proposes FSDETR, a…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Jianchao Huang , Fengming Zhang , Haibo Zhu , Tao Yan

We propose a new topological tool for computer vision - Scalar Function Topology Divergence (SFTD), which measures the dissimilarity of multi-scale topology between sublevel sets of two functions having a common domain. Functions can be…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Ilya Trofimov , Daria Voronkova , Eduard Tulchinskii , Evgeny Burnaev , Serguei Barannikov

With the rapid development of large multimodal models (LMMs), multimodal understanding applications are emerging. As most LMM inference requests originate from edge devices with limited computational capabilities, the predominant inference…

信号处理 · 电气工程与系统科学 2025-11-05 Cheng Yuan , Zhening Liu , Jiashu Lv , Jiawei Shao , Yufei Jiang , Jun Zhang , Xuelong Li

Indoor localization of autonomous mobile robots (AMRs) can be realized with fiducial markers. Such systems require only a simple, monocular camera as sensor and fiducial markers as passive, identifiable position references that can be…

图像与视频处理 · 电气工程与系统科学 2026-01-27 Sven Hinderer , Martina Scheffler , Bin Yang

Fiducial markers are a computer vision tool used for object pose estimation and detection. These markers are highly useful in fields such as industry, medicine and logistics. However, optimal lighting conditions are not always available,and…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Rafael Berral-Soler , Rafael Muñoz-Salinas , Rafael Medina-Carnicer , Manuel J. Marín-Jiménez

Light detection and ranging (LiDAR) systems are pivotal for precise distance and velocity measurement, yet widespread deployment requires solutions that balance their performance, robustness, and simplicity. Here, we propose a novel chaotic…

光学 · 物理学 2026-02-17 T. Wang , Z. Li , H. Shen , Y. Ma , Y. Li , S. Xiang , S. Baland , Y. Hao

This paper presents a real-time face detector, named Single Shot Scale-invariant Face Detector (S$^3$FD), which performs superiorly on various scales of faces with a single deep neural network, especially for small faces. Specifically, we…

计算机视觉与模式识别 · 计算机科学 2017-11-16 Shifeng Zhang , Xiangyu Zhu , Zhen Lei , Hailin Shi , Xiaobo Wang , Stan Z. Li

Globally consistent dense maps are a key requirement for long-term robot navigation in complex environments. While previous works have addressed the challenges of dense mapping and global consistency, most require more computational…

机器人学 · 计算机科学 2020-04-29 Victor Reijgwart , Alexander Millane , Helen Oleynikova , Roland Siegwart , Cesar Cadena , Juan Nieto

Visual localization is a fundamental task for various applications including autonomous driving and robotics. Prior methods focus on extracting large amounts of often redundant locally reliable features, resulting in limited efficiency and…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Fei Xue , Ignas Budvytis , Roberto Cipolla

We propose a fixed-lag smoother-based sensor fusion architecture to leverage the complementary benefits of range-based sensors and visual-inertial odometry (VIO) for localization. We use two fixed-lag smoothers (FLS) to decouple accurate…

机器人学 · 计算机科学 2024-01-05 Abhishek Goudar , Wenda Zhao , Angela P. Schoellig

Precise segmentation of retinal arteries and veins carries the diagnosis of systemic cardiovascular conditions. However, standard convolutional architectures often yield topologically disjointed segmentations, characterized by gaps and…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Iftekhar Ahmed , Shakib Absar , Aftar Ahmad Sami , Shadman Sakib , Debojyoti Biswas , Seraj Al Mahmud Mostafa

We introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-scale dataset, Mono3DRefer, which contains 3D object targets…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Yang Zhan , Yuan Yuan , Zhitong Xiong

3D lane detection and topology reasoning are essential tasks in autonomous driving scenarios, requiring not only detecting the accurate 3D coordinates on lane lines, but also reasoning the relationship between lanes and traffic elements.…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Han Li , Zehao Huang , Zitian Wang , Wenge Rong , Naiyan Wang , Si Liu

Vision-language pre-training (VLP) has emerged as an effective scheme for multimodal representation learning, but its reliance on large-scale multimodal data poses significant challenges for medical applications. Federated learning (FL)…

机器学习 · 计算机科学 2024-11-25 Zitao Shuai , Chenwei Wu , Zhengxu Tang , Liyue Shen