中文
相关论文

相关论文: Automated Segmentation and Tracking of Group House…

200 篇论文

Few-shot semantic segmentation has attracted growing interest for its ability to generalize to novel object categories using only a few annotated samples. To address data scarcity, recent methods incorporate multiple foundation models to…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Wei Zhuo , Zhiyue Tang , Wufeng Xue , Hao Ding , Junkai Ji , Linlin Shen

The primary challenge of cross-domain few-shot segmentation (CD-FSS) is the domain disparity between the training and inference phases, which can exist in either the input data or the target classes. Previous models struggle to learn…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Shi-Feng Peng , Guolei Sun , Yong Li , Hongsong Wang , Guo-Sen Xie

We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recognition,temporal localization, and spatiotemporal…

Fine-grained remote sensing image segmentation is essential for accurately identifying detailed objects in remote sensing images. Recently, vision transformer models (VTMs) pre-trained on large-scale datasets have demonstrated strong…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Shun Zhang , Xuechao Zou , Kai Li , Congyan Lang , Shiying Wang , Pin Tao , Tengfei Cao

While large-scale visual foundation models (VFMs) exhibit strong generalization across diverse visual domains, their potential for single-frame infrared small target (SIRST) detection remains largely unexplored. To fill this gap, we…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Chuang Yu , Jinmiao Zhao , Yunpeng Liu , Yaokun Li , Xiujun Shu , Yuanhao Feng , Bo Wang , Yimian Dai , Xiangyu Yue

Foundation Models (FMs) have shown impressive performance on various text and image processing tasks. They can generalize across domains and datasets in a zero-shot setting. This could make them suitable for automated quality inspection…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Simon Baeuerle , Pratik Khanna , Nils Friederich , Angelo Jovin Yamachui Sitcheu , Damir Shakirov , Andreas Steimer , Ralf Mikut

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

Endowing Large Multimodal Models (LMMs) with visual grounding capability can significantly enhance AIs' understanding of the visual world and their interaction with humans. However, existing methods typically fine-tune the parameters of…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Size Wu , Sheng Jin , Wenwei Zhang , Lumin Xu , Wentao Liu , Wei Li , Chen Change Loy

Very-High Resolution (VHR) remote sensing imagery is increasingly accessible, but often lacks annotations for effective machine learning applications. Recent foundation models like GroundingDINO and Segment Anything (SAM) provide…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Edoardo Arnaudo , Jacopo Lungo Vaschetti , Lorenzo Innocenti , Luca Barco , Davide Lisi , Vanina Fissore , Claudio Rossi

Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations to bridge the domain gap. However, these modalities face a…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Xuyang Chen , Conglang Zhang , Chuanheng Fu , Zihao Yang , Kaixuan Zhou , Yizhi Zhang , Jianan He , Yanfeng Zhang , Mingwei Sun , Zengmao Wang , Zhen Dong , Xiaoxiao Long , Liqiu Meng

Foundation models have shown strong performance in multi-object segmentation with visual prompts, yet histopathology images remain challenging due to high cellular density, heterogeneity, and the gap between pixel-level supervision and…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Yonghuang Wu , Wenwen Zeng , Xuan Xie , Chengqian Zhao , Guoqing Wu , Jinhua Yu

Medical image segmentation is crucial for clinical decision-making, but the scarcity of annotated data presents significant challenges. Few-shot segmentation (FSS) methods show promise but often require training on the target domain and…

计算机视觉与模式识别 · 计算机科学 2025-04-02 Lin Zhao , Xiao Chen , Eric Z. Chen , Yikang Liu , Terrence Chen , Shanhui Sun

Long-term visual localization is the problem of estimating the camera pose of a given query image in a scene whose appearance changes over time. It is an important problem in practice, for example, encountered in autonomous driving. In…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Måns Larsson , Erik Stenborg , Carl Toft , Lars Hammarstrand , Torsten Sattler , Fredrik Kahl

Following its success for vision and text, the "foundation model" (FM) paradigm -- pretraining large models on massive data, then fine-tuning on target tasks -- has rapidly expanded to domains in the sciences, engineering, healthcare, and…

机器学习 · 计算机科学 2025-03-24 Zongzhe Xu , Ritvik Gupta , Wenduo Cheng , Alexander Shen , Junhong Shen , Ameet Talwalkar , Mikhail Khodak

We present a zero-shot segmentation approach for agricultural imagery that leverages Plantnet, a large-scale plant classification model, in conjunction with its DinoV2 backbone and the Segment Anything Model (SAM). Rather than collecting…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Simon Ravé , Jean-Christophe Lombardo , Pejman Rasti , Alexis Joly , David Rousseau

Image segmentation foundation models (SFMs) like Segment Anything Model (SAM) have achieved impressive zero-shot and interactive segmentation across diverse domains. However, they struggle to segment objects with certain structures,…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Yixin Zhang , Nicholas Konz , Kevin Kramer , Maciej A. Mazurowski

Foundation segmentation models such as Segment Anything Model (SAM) are now routinely used in iterative pipelines, where each predicted mask is fed back as the next prompt. This practice turns segmentation into a closed-loop dynamical…

计算机视觉与模式识别 · 计算机科学 2026-05-26 H. M. Shadman Tabib , Md. Shamsuzzoha Bayzid , M Sohel Rahman

Precise identification of individual cows is a fundamental prerequisite for comprehensive digital management in smart livestock farming. While existing animal identification methods excel in controlled, single-camera settings, they face…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Runcheng Wang , Yaru Chen , Guiguo Zhang , Honghua Jiang , Yongliang Qiao

DepthCropSeg++: a foundation model for crop segmentation, capable of segmenting different crop species under open in-field environment. Crop segmentation is a fundamental task for modern agriculture, which closely relates to many downstream…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Jiafei Zhang , Songliang Cao , Binghui Xu , Yanan Li , Weiwei Jia , Tingting Wu , Hao Lu , Weijuan Hu , Zhiguo Han

Foundation models (FMs) are a popular topic of research in AI. Their ability to generalize to new tasks and datasets without retraining or needing an abundance of data makes them an appealing candidate for applications on specialist…

计算机视觉与模式识别 · 计算机科学 2024-09-06 Marga Don , Stijn Pinson , Blanca Guillen Cebrian , Yuki M. Asano