中文
相关论文

相关论文: TinyViT: Field Deployable Transformer Pipeline for…

200 篇论文

Reliable long-term deployment of autonomous robots in agricultural environments remains challenging due to perceptual aliasing, seasonal variability, and the dynamic nature of crop canopies. Vineyards, characterized by repetitive row…

机器人学 · 计算机科学 2026-03-06 Giorgio Audrito , Mauro Martini , Alessandro Navone , Giorgia Galluzzo , Marcello Chiaberge

Detecting unseen anomalies in unstructured environments presents a critical challenge for industrial and agricultural applications such as material recycling and weeding. Existing perception systems frequently fail to satisfy the strict…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Melanie Neubauer , Elmar Rueckert , Christian Rauch

Modern microscopy routinely produces gigapixel images that contain structures across multiple spatial scales, from fine cellular morphology to broader tissue organization. Many analysis tasks require combining these scales, yet most vision…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Albert Dominguez Mantes , Gioele La Manno , Martin Weigert

Vision Transformers (ViTs) have revolutionized computer vision by leveraging self-attention to model long-range dependencies. However, ViTs face challenges such as high computational costs due to the quadratic scaling of self-attention and…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Zhoujie Qian

Vision Transformers (ViTs) are built by stacking independently parameterized blocks, but it remains unclear how much of this depth requires layer specific transformations and how much can be realized through recurrent computation. We study…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Michal Byra , Pawel Olszowiec , Grzegorz Stefanski , Grzegorz Gruszczynski , Alberto Presta

Vision transformers have emerged as a promising alternative to convolutional neural networks for various image analysis tasks, offering comparable or superior performance. However, one significant drawback of ViTs is their…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Kaixin Xu , Zhe Wang , Chunyun Chen , Xue Geng , Jie Lin , Mohamed M. Sabry Aly , Xulei Yang , Min Wu , Xiaoli Li , Weisi Lin

Soiling is the accumulation of dirt in solar panels which leads to a decreasing trend in solar energy yield and may be the cause of vast revenue losses. The effect of soiling can be reduced by washing the panels, which is, however, a…

信号处理 · 电气工程与系统科学 2023-01-31 Alexandros Kalimeris , Ioannis Psarros , Giorgos Giannopoulos , Manolis Terrovitis , George Papastefanatos , Gregory Kotsis

We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing imagery, aiming to support and enhance the Emergent Value Added Product (EVAP) developed by the Taiwan…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yi-Shan Chu , Hsuan-Cheng Wei

Transient photovoltage (TPV) is a technique frequently used to determine charge carrier lifetimes in thin-film solar cells such as organic, dye sensitized and perovskite solar cells. As this lifetime is often incident light intensity…

应用物理 · 物理学 2019-08-06 Oskar J. Sandberg , Kristofer Tvingstedt , Paul Meredith , Ardalan Armin

In this paper we introduce the Temporo-Spatial Vision Transformer (TSViT), a fully-attentional model for general Satellite Image Time Series (SITS) processing based on the Vision Transformer (ViT). TSViT splits a SITS record into…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Michail Tarasiou , Erik Chavez , Stefanos Zafeiriou

Accurate and timely identification of plant leaf diseases is essential for resilient and sustainable agriculture, yet most deep learning approaches rely on large annotated datasets and computationally intensive models that are unsuitable…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Anika Islam , Tasfia Tahsin , Zaarin Anjum , Md. Bakhtiar Hasan , Md. Hasanul Kabir

Skin lesion segmentation (SLS) plays an important role in skin lesion analysis. Vision transformers (ViTs) are considered an auspicious solution for SLS, but they require more training data compared to convolutional neural networks (CNNs)…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Siyi Du , Nourhan Bayasi , Ghassan Hamarneh , Rafeef Garbi

The increase in the use of photovoltaic (PV) energy in the world has shown that the useful life and maintenance of a PV plant directly depend on theability to quickly detect severe faults on a PV plant. To solve this problem of detection,…

Vision Transformers (ViT) have recently demonstrated success across a myriad of computer vision tasks. However, their elevated computational demands pose significant challenges for real-world deployment. While low-rank approximation stands…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Chi-Chih Chang , Yuan-Yao Sung , Shixing Yu , Ning-Chi Huang , Diana Marculescu , Kai-Chiang Wu

Land-cover underpins ecosystem services, hydrologic regulation, disaster-risk reduction, and evidence-based land planning; timely, accurate land-cover maps are therefore critical for environmental stewardship. Remote sensing-based…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Kai Wang , Siyi Chen , Weicong Pang , Chenchen Zhang , Renjun Gao , Ziru Chen , Cheng Li , Dasa Gu , Rui Huang , Alexis Kai Hon Lau

In this study, we introduce an enhanced version of ViT that conducts attention-based QKV operations during the initial stages of downsampling. Performing attention directly on high-resolution feature maps is computationally demanding due to…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Bohang Sun

Background and objective: Cell-level pathological image analysis requires working with extremely small image patches (40x40 pixels), far below standard ImageNet resolutions. It remains unclear whether modern deep learning architectures and…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Hiroki Kagiyama , Toru Nagasaka , Yukari Adachi , Takaaki Tachibana , Ryota Ito , Mitsugu Fujita , Kimihiro Yamashita , Yoshihiro Kakeji

One of the major goals of tomorrow's agriculture is to increase agricultural productivity but above all the quality of production while significantly reducing the use of inputs. Meeting this goal is a real scientific and technological…

图像与视频处理 · 电气工程与系统科学 2020-05-14 Mohamed Kerkech , Adel Hafiane , Raphael Canals

Vision Transformers (ViTs) are essential as foundation backbones in establishing the visual comprehension capabilities of Multimodal Large Language Models (MLLMs). Although most ViTs achieve impressive performance through image-text…

计算机视觉与模式识别 · 计算机科学 2025-07-03 Weijie Yin , Dingkang Yang , Hongyuan Dong , Zijian Kang , Jiacong Wang , Xiao Liang , Chao Feng , Jiao Ran

Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is severely hindered by the computational complexity and global reduction bottleneck imposed by layer…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Kieran Carrigg , Sigur de Vries , Amirhossein Sadough , Marcel van Gerven