中文
相关论文

相关论文: FSDAM: Few-Shot Driving Attention Modeling via Vis…

200 篇论文

Vision-Language Models (vLLMs) have emerged as powerful architectures for joint reasoning over visual and textual inputs, enabling breakthroughs in image captioning, cross modal retrieval, and multimodal dialogue. However, as these models…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Andrew Kiruluta , Preethi Raju , Priscilla Burity

EASA's learning-assurance guidance requires data-driven aviation systems to build and monitor their own situation representation, yet for neural networks the technical means to provide such evidence remain an open problem. We address this…

机器学习 · 计算机科学 2026-05-21 Romeo Valentin , Olivia Beyer Bruvik , Marc R. Schlichting , Mykel J. Kochenderfer

Few-shot learning (FSL) has attracted considerable attention recently. Among existing approaches, the metric-based method aims to train an embedding network that can make similar samples close while dissimilar samples as far as possible and…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Bin Xiao , Chien-Liang Liu , Wen-Hoar Hsaio

Few-shot video object segmentation (FSVOS) aims to segment dynamic objects of unseen classes by resorting to a small set of support images that contain pixel-level object annotations. Existing methods have demonstrated that the domain…

计算机视觉与模式识别 · 计算机科学 2023-07-18 Yin Tang , Tao Chen , Xiruo Jiang , Yazhou Yao , Guo-Sen Xie , Heng-Tao Shen

Few-Shot Semantic Segmentation (FSS) models achieve strong performance in segmenting novel classes with minimal labeled examples, yet their decision-making processes remain largely opaque. While explainable AI has advanced significantly in…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Pasquale De Marinis , Uzay Kaymak , Rogier Brussee , Gennaro Vessio , Giovanna Castellano

Driver distraction detection is essential for improving traffic safety and reducing road accidents. However, existing models often suffer from degraded generalization when deployed in real-world scenarios. This limitation primarily arises…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Haibin Sun , Xinghui Song

In this work, we propose Few Shot Domain Adapting Graph (FS-DAG), a scalable and efficient model architecture for visually rich document understanding (VRDU) in few-shot settings. FS-DAG leverages domain-specific and language/vision…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Amit Agarwal , Srikant Panda , Kulbhushan Pachauri

Purpose: Surgical workflow analysis is crucial for improving surgical efficiency and safety. However, previous studies rely heavily on large-scale annotated datasets, posing challenges in cost, scalability, and reliance on expert…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Tingxuan Chen , Kun Yuan , Vinkle Srivastav , Nassir Navab , Nicolas Padoy

Robust and efficient learning remains a challenging problem in robotics, in particular with complex visual inputs. Inspired by human attention mechanism, with which we quickly process complex visual scenes and react to changes in the…

机器人学 · 计算机科学 2023-08-30 Daniel Scheuchenstuhl , Stefan Ulmer , Felix Resch , Luigi Berducci , Radu Grosu

Cross-view object geo-localization has recently gained attention due to potential applications. Existing methods aim to capture spatial dependencies of query objects between different views through attention mechanisms to obtain spatial…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Xingtao Ling Yingying Zhu

Fine-grained truck classification is critical for intelligent transportation systems (ITS), yet current LiDAR-based methods face scalability challenges due to their reliance on supervised deep learning and labor-intensive manual annotation.…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Yiqiao Li , Bo Shang , Jie Wei

Few-shot anomaly detection (FSAD) has made significant strides, yet existing methods still face critical challenges: (i) dependence on task- or dataset-specific training/fine-tuning, (ii) reliance on language supervision or carefully…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Guohuan Xie , Xin He , Dingying Fan , Siqi Li , Yun Liu

Recognizing discriminative details such as eyes and beaks is important for distinguishing fine-grained classes since they have similar overall appearances. In this regard, we introduce Task Discrepancy Maximization (TDM), a simple module…

计算机视觉与模式识别 · 计算机科学 2022-07-06 SuBeen Lee , WonJun Moon , Jae-Pil Heo

Few-shot learning presents a critical solution for cancer diagnosis in computational pathology (CPath), addressing fundamental limitations in data availability, particularly the scarcity of expert annotations and patient privacy…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Zhengrui Guo , Conghao Xiong , Jiabo Ma , Qichen Sun , Lishuang Feng , Jinzhuo Wang , Hao Chen

Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Jinxiang Lai , Siqian Yang , Wenlong Wu , Tao Wu , Guannan Jiang , Xi Wang , Jun Liu , Bin-Bin Gao , Wei Zhang , Yuan Xie , Chengjie Wang

Recently, masked image modeling (MIM), which learns visual representations by reconstructing the masked patches of an image, has dominated self-supervised learning in computer vision. However, the pre-training of MIM always takes massive…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Jie Gui , Tuo Chen , Minjing Dong , Zhengqi Liu , Hao Luo , James Tin-Yau Kwok , Yuan Yan Tang

Vision-Language-Action (VLA) models offer significant potential for end-to-end driving, yet their reasoning is often constrained by textual Chains-of-Thought (CoT). This symbolic compression of visual information creates a modality gap…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Shuang Zeng , Xinyuan Chang , Mengwei Xie , Xinran Liu , Yifan Bai , Zheng Pan , Mu Xu , Xing Wei , Ning Guo

Intention prediction is a crucial task for Autonomous Driving (AD). Due to the variety of size and layout of intersections, it is challenging to predict intention of human driver at different intersections, especially unseen and irregular…

机器人学 · 计算机科学 2021-03-10 Fei Li , Xiangxu Li , Jun Luo , Shiwei Fan , Hongbo Zhang

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and image regions. This…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Mir Rayat Imtiaz Hossain , Mennatullah Siam , Leonid Sigal , James J. Little

In recent years, deep learning has been widely applied in communications and achieved remarkable performance improvement. Most of the existing works are based on data-driven deep learning, which requires a significant amount of training…

信息论 · 计算机科学 2022-09-07 Ouya Wang , Jiabao Gao , Geoffrey Ye Li