English
Related papers

Related papers: Adaptive Subsampling for ROI-based Visual Tracking…

200 papers

Affordance detection aims to jointly address the fundamental "what-where-how" challenge in embodied AI by understanding "what" an object is, "where" the object is located, and "how" it can be used. However, most affordance learning methods…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Yuqi Ji , Junjie Ke , Lihuo He , Jun Liu , Kaifan Zhang , Yu-Kun Lai , Guiguang Ding , Xinbo Gao

Latest deep learning methods for object detection provide remarkable performance, but have limits when used in robotic applications. One of the most relevant issues is the long training time, which is due to the large size and imbalance of…

Robotics · Computer Science 2021-06-30 Elisa Maiettini , Giulia Pasquale , Lorenzo Rosasco , Lorenzo Natale

Optical satellite-to-ground communication (OSGC) has the potential to improve access to fast and affordable Internet in remote regions. Atmospheric turbulence, however, distorts the optical beam, eroding the data rate potential when…

Machine Learning · Computer Science 2023-03-15 Payam Parvizi , Runnan Zou , Colin Bellinger , Ross Cheriton , Davide Spinello

Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on static dense computation that applies uniform processing to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Xu Lin , Jinlong Peng , Zhenye Gan , Jiawen Zhu , Jun Liu

Vision-Language Models (VLMs) have demonstrated strong performance on multimodal reasoning tasks, but their deployment remains challenging due to high inference latency and computational cost, particularly when processing high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Putu Indah Githa Cahyani , Komang David Dananjaya Suartana , Novanto Yudistira

We propose a camera-based assistive text reading framework to help blind persons read text labels and product packaging from hand-held objects in their daily life. To isolate the object from untidy backgrounds or other surrounding objects…

Human-Computer Interaction · Computer Science 2019-01-18 Rajkumar N , Anand M. G , Barathiraja N

Downsampling is widely adopted to achieve a good trade-off between accuracy and latency for visual recognition. Unfortunately, the commonly used pooling layers are not learned, and thus cannot preserve important information. As another…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Ho Man Kwan , Shenghui Song

Recently, foundation models based on Vision Transformers (ViTs) have become widely available. However, their fine-tuning process is highly resource-intensive, and it hinders their adoption in several edge or low-energy applications. To this…

Computer Vision and Pattern Recognition · Computer Science 2024-08-19 Alessio Devoto , Federico Alvetreti , Jary Pomponi , Paolo Di Lorenzo , Pasquale Minervini , Simone Scardapane

Aggregating deep convolutional features into a global image vector has attracted sustained attention in image retrieval. In this paper, we propose an efficient unsupervised aggregation method that uses an adaptive Gaussian filter and an…

Computer Vision and Pattern Recognition · Computer Science 2018-03-21 Jiaxing Wang , Jihua Zhu , Shanmin Pang , Zhongyu Li , Yaochen Li , Xueming Qian

Vision algorithms can be executed directly on the image sensor when implemented on the next-generation sensors known as focal-plane sensor-processor arrays (FPSP)s, where every pixel has a processor. FPSPs greatly improve latency, reducing…

Robotics · Computer Science 2025-10-07 Matthew Lisondra , Junseo Kim , Glenn Takashi Shimoda , Kourosh Zareinia , Sajad Saeedi

High-precision surface defect detection in manufacturing is essential for ensuring quality control. Laser triangulation profilometric sensors are key to this process, providing detailed and accurate surface measurements over a line. To…

Robotics · Computer Science 2024-09-06 Sara Roos-Hoefgeest , Mario Roos-Hoefgeest , Ignacio Alvarez , Rafael C. González

While traditional Deep Learning (DL) optimization methods treat all training samples equally, Distributionally Robust Optimization (DRO) adaptively assigns importance weights to different samples. However, a significant gap exists between…

This paper investigates how to achieve both low-power operations of sensor nodes and accurate state estimation using Kalman filter for internet of things (IoT) monitoring employing wireless sensor networks under radio resource constraint.…

Networking and Internet Architecture · Computer Science 2025-07-22 Takaho Shimokasa , Hiroyuki Yomo , Federico Chiariotti , Junya Shiraishi , Petar Popovski

Medical imaging systems are commonly assessed and optimized by the use of objective measures of image quality (IQ). The performance of the ideal observer (IO) acting on imaging measurements has long been advocated as a figure-of-merit to…

Medical Physics · Physics 2025-01-17 Kaiyan Li , Prabhat Kc , Hua Li , Kyle J. Myers , Mark A. Anastasio , Rongping Zeng

Advances in deep vision techniques and ubiquity of smart cameras will drive the next generation of video analytics. However, video analytics applications consume vast amounts of energy as both deep learning techniques and cameras are…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Yoones Rezaei , Stephen Lee , Daniel Mosse

Inter-subject registration of cortical areas is necessary in functional imaging (fMRI) studies for making inferences about equivalent brain function across a population. However, many high-level visual brain areas are defined as peaks of…

Neurons and Cognition · Quantitative Biology 2016-06-09 Marius Cătălin Iordan , Armand Joulin , Diane M. Beck , Li Fei-Fei

Visual Simultaneous Localisation and Mapping (VSLAM) is a key enabling technology for small embedded robotic systems such as aerial vehicles. Recent advances in equivariant filter and observer design offer the potential of a new generation…

Robotics · Computer Science 2020-06-01 Pieter van Goor , Robert Mahony , Tarek Hamel , Jochen Trumpf

Eye-tracking technology is integral to numerous consumer electronics applications, particularly in the realm of virtual and augmented reality (VR/AR). These applications demand solutions that excel in three crucial aspects: low-latency,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Baoheng Zhang , Yizhao Gao , Jingyuan Li , Hayden Kwok-Hay So

Motivated by the Parameter-Efficient Fine-Tuning (PEFT) in large language models, we propose LoRAT, a method that unveils the power of large ViT model for tracking within laboratory-level resources. The essence of our work lies in adapting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Liting Lin , Heng Fan , Zhipeng Zhang , Yaowei Wang , Yong Xu , Haibin Ling

Foundation models for vision are predominantly trained on RGB data, while many safety-critical applications rely on non-visible modalities such as infrared (IR) and synthetic aperture radar (SAR). We study whether a single flow-matching…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Maxim Clouser , Kia Khezeli , John Kalantari