中文
相关论文

相关论文: MobileGeo: Exploring Hierarchical Knowledge Distil…

200 篇论文

Object detection using images or videos captured by drones is a promising technology with significant potential across various industries. However, a major challenge is that drone images are typically taken from high altitudes, making…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Hyun-Ki Jung

The empowering unmanned aerial vehicles (UAVs) have been extensively used in providing intelligence such as target tracking. In our field experiments, a pre-trained convolutional neural network (CNN) is deployed at the UAV to identify a…

图像与视频处理 · 电气工程与系统科学 2020-08-19 Bo Yang , Xuelin Cao , Chau Yuen , Lijun Qian

Multi-view data capture permits free-viewpoint video (FVV) content creation. To this end, several users must capture video streams, calibrated in both time and pose, framing the same object/scene, from different viewpoints. New-generation…

多媒体 · 计算机科学 2020-05-08 Matteo Bortolon , Paul Chippendale , Stefano Messelodi , Fabio Poiesi

Recent advances in machine learning and hardware have produced embedded devices capable of performing real-time object detection with commendable accuracy. We consider a scenario in which embedded devices rely on an onboard object detector,…

分布式、并行与集群计算 · 计算机科学 2024-10-25 Jiaming Qiu , Ruiqi Wang , Brooks Hu , Roch Guerin , Chenyang Lu

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introduce DiG (Differential…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Zhou Tao , Shida Wang , Yongxiang Hua , Haoyu Cao , Linli Xu

The deployment of unmanned aerial vehicles (UAVs) in many different settings has provided various solutions and strategies for networking paradigms. Therefore, it reduces the complexity of the developments for the existing problems, which…

网络与互联网体系结构 · 计算机科学 2025-02-25 Baris Yamansavascilar , Atay Ozgovde , Cem Ersoy

Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Siyu Zhang , Lianlei Shan , Runhe Qiu

Object detection and classification using video is necessary for intelligent planning and navigation on a mobile robot. However, current methods can be too slow or not sufficient for distinguishing multiple classes. Techniques that rely on…

计算机视觉与模式识别 · 计算机科学 2011-11-08 Colin S. Lea , Jason J. Corso

Due to limited resources on edge and different characteristics of deep neural network (DNN) models, it is a big challenge to optimize DNN inference performance in terms of energy consumption and end-to-end latency on edge devices. In…

机器学习 · 计算机科学 2023-06-26 Ziyang Zhang , Yang Zhao , Huan Li , Changyao Lin , Jie Liu

Deploying deep neural networks on mobile devices is increasingly important but remains challenging due to limited computing resources. On the other hand, their unified memory architecture and narrower gap between CPU and GPU performance…

机器学习 · 计算机科学 2026-02-20 Zhuojin Li , Marco Paolieri , Leana Golubchik

Multimodal large language model (MLLM) inference splits into two phases with opposing hardware demands: vision encoding is compute-bound, while language generation is memory-bandwidth-bound. We show that under standard transformer KV…

机器学习 · 计算机科学 2026-03-16 Donglin Yu

Object detection techniques that achieve state-of-the-art detection accuracy employ convolutional neural networks, implemented to have optimal performance in graphics processing units. Some hardware systems, such as mobile robots, operate…

In image processing, it is essential to detect and track air targets, especially UAVs. In this paper, we detect the flying drone using a fisheye camera. In the field of diagnosis and classification of objects, there are always many problems…

计算机视觉与模式识别 · 计算机科学 2021-07-06 Fatemeh Mahdavi , Roozbeh Rajabi

Most existing cross-view object geo-localization approaches adopt anchor-based paradigm. Although effective, such methods are inherently constrained by predefined anchors. To eliminate this dependency, we first propose an anchor-free…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Xingtao Ling , Chenlin Fu , Yingying Zhu

Herbage mass yield and composition estimation is an important tool for dairy farmers to ensure an adequate supply of high quality herbage for grazing and subsequently milk production. By accurately estimating herbage mass and composition,…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Paul Albert , Mohamed Saadeldin , Badri Narayanan , Jaime Fernandez , Brian Mac Namee , Deirdre Hennessey , Noel E. O'Connor , Kevin McGuinness

Unmanned Aerial Vehicles (UAVs) especially drones, equipped with vision techniques have become very popular in recent years, with their extensive use in wide range of applications. Many of these applications require use of computer vision…

计算机视觉与模式识别 · 计算机科学 2019-06-04 Subrahmanyam Vaddi , Chandan Kumar , Ali Jannesari

Live tracking of wildlife via high-resolution video processing directly onboard drones is widely unexplored and most existing solutions rely on streaming video to ground stations to support navigation. Yet, both autonomous animal-reactive…

计算机视觉与模式识别 · 计算机科学 2025-05-26 Nguyen Ngoc Dat , Tom Richardson , Matthew Watson , Kilian Meier , Jenna Kline , Sid Reid , Guy Maalouf , Duncan Hine , Majid Mirmehdi , Tilo Burghardt

Augmented Reality (AR) and Virtual Reality (VR) systems involve computationally intensive image processing algorithms that can burden end-devices with limited resources, leading to poor performance in providing low latency services.…

分布式、并行与集群计算 · 计算机科学 2024-03-20 Mohammadsadeq Garshasbi Herabad , Javid Taheri , Bestoun S. Ahmed , Calin Curescu

Gaze estimation, the task of predicting where an individual is looking, is a critical task with direct applications in areas such as human-computer interaction and virtual reality. Estimating the direction of looking in unconstrained…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Andy Cătrună , Adrian Cosma , Emilian Rădoi

Cross-view object geo-localization (CVOGL) aims to determine the location of a specific object in high-resolution satellite imagery given a query image with a point prompt. Existing approaches treat CVOGL as a one-shot detection task,…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Xiaohan Zhang , Si-Yuan Cao , Xiaokai Bai , Yiming Li , Zhangkai Shen , Zhe Wu , Xiaoxi Hu , Hui-liang Shen