中文
相关论文

相关论文: PiCo: Active Manifold Canonicalization for Robust …

200 篇论文

Visually impaired users face significant challenges in daily information access and real-time environmental perception, and there is an urgent need for intelligent assistive systems with accurate recognition capabilities. Although…

Recent advances in vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities, yet adapting these models to specialized domains remains a significant challenge. Building on recent theoretical insights suggesting that…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Pranav Mantini , Shishir K. Shah

Recent advances in multimodal models have demonstrated impressive capabilities in object recognition and scene understanding. However, these models often struggle with precise spatial localization - a critical capability for real-world…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Joongwon Chae , Zhenyu Wang , Lian Zhang , Dongmei Yu , Peiwu Qin

This paper presents ViTOC (Vision Transformer and Object-aware Captioner), a novel vision-language model for image captioning that addresses the challenges of accuracy and diversity in generated descriptions. Unlike conventional approaches,…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Feiyang Huang

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

The capacity to recognize faces under varied poses is a fundamental human ability that presents a unique challenge for computer vision systems. Compared to frontal face recognition, which has been intensively studied and has gradually…

计算机视觉与模式识别 · 计算机科学 2016-07-19 Changxing Ding , Dacheng Tao

Explicit communication is often valued for its directness in presenting information but requires attention during exchange, resulting in cognitive interruptions. On the other hand, implicit communication contributes to tacit and smooth…

机器人学 · 计算机科学 2025-10-02 Andrew Boateng , Prakhar Bhartiya , Taha Shaheen , Yu Zhang

The history of computing started with analog computers consisting of physical devices performing specialized functions such as predicting the trajectory of cannon balls. In modern times, this idea has been extended, for example, to…

图像与视频处理 · 电气工程与系统科学 2022-08-29 Callen MacPhee , Bahram Jalali

Odometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient…

计算机视觉与模式识别 · 计算机科学 2020-03-06 Wei Wang , Muhamad Risqi U. Saputra , Peijun Zhao , Pedro Gusmao , Bo Yang , Changhao Chen , Andrew Markham , Niki Trigoni

In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physically incorrect. For instance, when estimating the poses and shapes of objects in the scene…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Xihang Yu , Rajat Talak , Lorenzo Shaikewitz , Luca Carlone

Occlusion-aware decision-making is essential in autonomous driving due to the high uncertainty of various occlusions. Recent occlusion-aware decision-making methods encounter issues such as high computational complexity, scenario…

机器人学 · 计算机科学 2025-04-10 Jie Jia , Yiming Shu , Zhongxue Gan , Wenchao Ding

Visual SLAM approaches typically depend on loop closure detection to correct the inconsistencies that may arise during the map and camera trajectory calculations, typically making use of point features for detecting and closing the existing…

计算机视觉与模式识别 · 计算机科学 2020-09-22 Joan P. Company-Corcoles , Emilio Garcia-Fidalgo , Alberto Ortiz

The pose problem is one of the bottlenecks in automatic face recognition. We argue that one of the diffculties in this problem is the severe misalignment in face images or feature vectors with different poses. In this paper, we propose that…

计算机视觉与模式识别 · 计算机科学 2015-07-30 Annan Li , Shiguang Shan , Xilin Chen , Bingpeng Ma , Shuicheng Yan , Wen Gao

The success of machine learning for real-world robotic systems has created a new form of intellectual property: the trained policy. This raises a critical need for novel methods that verify ownership and detect unauthorized, possibly unsafe…

机器人学 · 计算机科学 2025-12-18 Michael Amir , Manon Flageat , Amanda Prorok

Oculomotor alterations constitute a promising biomarker to detect and characterize Parkinson's disease (PD), even in prodromal stages. Currently, only global and simplified eye movement trajectories are employed to approximate the complex…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Juan Niño , Luis Guayacán , Santiago Gómez , Fabio Martínez

Monocular 3D object detection typically relies on pseudo-labeling techniques to reduce dependency on real-world annotations. Recent advances demonstrate that deterministic linguistic cues can serve as effective auxiliary weak supervision…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Chupeng Liu , Jiyong Rao , Shangquan Sun , Runkai Zhao , Weidong Cai

Video Anomaly Detection (VAD) has traditionally been framed as binary classification or outlier detection, providing neither interpretable reasoning nor precise spatial localization of anomalous events. While Vision-Language Models (VLMs)…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Sakshi Agarwal , Aishik Konwer , Ankit Parag Shah

Learning model-free object pose estimation for unseen instances remains a fundamental challenge in 3D vision. Existing methods typically fall into two disjoint paradigms: category-level approaches predict absolute poses in a canonical space…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Weihang Li , Lorenzo Garattoni , Fabien Despinoy , Nassir Navab , Benjamin Busam

6D pose estimation of textureless shiny objects has become an essential problem in many robotic applications. Many pose estimators require high-quality depth data, often measured by structured light cameras. However, when objects have shiny…

机器人学 · 计算机科学 2023-08-29 Jun Yang , Jian Yao , Steven L. Waslander

Visual Anomaly Detection (VAD) seeks to identify abnormal images and precisely localize the corresponding anomalous regions, relying solely on normal data during training. This approach has proven essential in domains such as manufacturing…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Manuel Barusco , Francesco Borsatti , Nicola Beda , Davide Dalle Pezze , Gian Antonio Susto