中文
相关论文

相关论文: Adaptive Subsampling for ROI-based Visual Tracking…

200 篇论文

Affordance detection aims to jointly address the fundamental "what-where-how" challenge in embodied AI by understanding "what" an object is, "where" the object is located, and "how" it can be used. However, most affordance learning methods…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Yuqi Ji , Junjie Ke , Lihuo He , Jun Liu , Kaifan Zhang , Yu-Kun Lai , Guiguang Ding , Xinbo Gao

Latest deep learning methods for object detection provide remarkable performance, but have limits when used in robotic applications. One of the most relevant issues is the long training time, which is due to the large size and imbalance of…

机器人学 · 计算机科学 2021-06-30 Elisa Maiettini , Giulia Pasquale , Lorenzo Rosasco , Lorenzo Natale

Optical satellite-to-ground communication (OSGC) has the potential to improve access to fast and affordable Internet in remote regions. Atmospheric turbulence, however, distorts the optical beam, eroding the data rate potential when…

机器学习 · 计算机科学 2023-03-15 Payam Parvizi , Runnan Zou , Colin Bellinger , Ross Cheriton , Davide Spinello

Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on static dense computation that applies uniform processing to…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Xu Lin , Jinlong Peng , Zhenye Gan , Jiawen Zhu , Jun Liu

Vision-Language Models (VLMs) have demonstrated strong performance on multimodal reasoning tasks, but their deployment remains challenging due to high inference latency and computational cost, particularly when processing high-resolution…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Putu Indah Githa Cahyani , Komang David Dananjaya Suartana , Novanto Yudistira

We propose a camera-based assistive text reading framework to help blind persons read text labels and product packaging from hand-held objects in their daily life. To isolate the object from untidy backgrounds or other surrounding objects…

人机交互 · 计算机科学 2019-01-18 Rajkumar N , Anand M. G , Barathiraja N

Downsampling is widely adopted to achieve a good trade-off between accuracy and latency for visual recognition. Unfortunately, the commonly used pooling layers are not learned, and thus cannot preserve important information. As another…

计算机视觉与模式识别 · 计算机科学 2022-07-26 Ho Man Kwan , Shenghui Song

Recently, foundation models based on Vision Transformers (ViTs) have become widely available. However, their fine-tuning process is highly resource-intensive, and it hinders their adoption in several edge or low-energy applications. To this…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Alessio Devoto , Federico Alvetreti , Jary Pomponi , Paolo Di Lorenzo , Pasquale Minervini , Simone Scardapane

Aggregating deep convolutional features into a global image vector has attracted sustained attention in image retrieval. In this paper, we propose an efficient unsupervised aggregation method that uses an adaptive Gaussian filter and an…

计算机视觉与模式识别 · 计算机科学 2018-03-21 Jiaxing Wang , Jihua Zhu , Shanmin Pang , Zhongyu Li , Yaochen Li , Xueming Qian

Vision algorithms can be executed directly on the image sensor when implemented on the next-generation sensors known as focal-plane sensor-processor arrays (FPSP)s, where every pixel has a processor. FPSPs greatly improve latency, reducing…

机器人学 · 计算机科学 2025-10-07 Matthew Lisondra , Junseo Kim , Glenn Takashi Shimoda , Kourosh Zareinia , Sajad Saeedi

High-precision surface defect detection in manufacturing is essential for ensuring quality control. Laser triangulation profilometric sensors are key to this process, providing detailed and accurate surface measurements over a line. To…

机器人学 · 计算机科学 2024-09-06 Sara Roos-Hoefgeest , Mario Roos-Hoefgeest , Ignacio Alvarez , Rafael C. González

While traditional Deep Learning (DL) optimization methods treat all training samples equally, Distributionally Robust Optimization (DRO) adaptively assigns importance weights to different samples. However, a significant gap exists between…

This paper investigates how to achieve both low-power operations of sensor nodes and accurate state estimation using Kalman filter for internet of things (IoT) monitoring employing wireless sensor networks under radio resource constraint.…

网络与互联网体系结构 · 计算机科学 2025-07-22 Takaho Shimokasa , Hiroyuki Yomo , Federico Chiariotti , Junya Shiraishi , Petar Popovski

Medical imaging systems are commonly assessed and optimized by the use of objective measures of image quality (IQ). The performance of the ideal observer (IO) acting on imaging measurements has long been advocated as a figure-of-merit to…

医学物理 · 物理学 2025-01-17 Kaiyan Li , Prabhat Kc , Hua Li , Kyle J. Myers , Mark A. Anastasio , Rongping Zeng

Advances in deep vision techniques and ubiquity of smart cameras will drive the next generation of video analytics. However, video analytics applications consume vast amounts of energy as both deep learning techniques and cameras are…

计算机视觉与模式识别 · 计算机科学 2022-04-22 Yoones Rezaei , Stephen Lee , Daniel Mosse

Inter-subject registration of cortical areas is necessary in functional imaging (fMRI) studies for making inferences about equivalent brain function across a population. However, many high-level visual brain areas are defined as peaks of…

神经元与认知 · 定量生物学 2016-06-09 Marius Cătălin Iordan , Armand Joulin , Diane M. Beck , Li Fei-Fei

Visual Simultaneous Localisation and Mapping (VSLAM) is a key enabling technology for small embedded robotic systems such as aerial vehicles. Recent advances in equivariant filter and observer design offer the potential of a new generation…

机器人学 · 计算机科学 2020-06-01 Pieter van Goor , Robert Mahony , Tarek Hamel , Jochen Trumpf

Eye-tracking technology is integral to numerous consumer electronics applications, particularly in the realm of virtual and augmented reality (VR/AR). These applications demand solutions that excel in three crucial aspects: low-latency,…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Baoheng Zhang , Yizhao Gao , Jingyuan Li , Hayden Kwok-Hay So

Motivated by the Parameter-Efficient Fine-Tuning (PEFT) in large language models, we propose LoRAT, a method that unveils the power of large ViT model for tracking within laboratory-level resources. The essence of our work lies in adapting…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Liting Lin , Heng Fan , Zhipeng Zhang , Yaowei Wang , Yong Xu , Haibin Ling

Foundation models for vision are predominantly trained on RGB data, while many safety-critical applications rely on non-visible modalities such as infrared (IR) and synthetic aperture radar (SAR). We study whether a single flow-matching…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Maxim Clouser , Kia Khezeli , John Kalantari