中文
相关论文

相关论文: MMLA: Multi-Environment, Multi-Species, Low-Altitu…

200 篇论文

Perception of Low-Altitude Aircraft (LAA) in 3D space enables precise 3D object localization and behavior understanding. However, datasets tailored for 3D LAA perception remain scarce. To address this gap, we present LAA3D, a large-scale…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Hai Wu , Shuai Tang , Jiale Wang , Longkun Zou , Mingyue Guo , Rongqin Liang , Ke Chen , Yaowei Wang

Open-vocabulary detection (OVD) aims to detect objects beyond a predefined set of categories. As a pioneering model incorporating the YOLO series into OVD, YOLO-World is well-suited for scenarios prioritizing speed and efficiency. However,…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Haoxuan Wang , Qingdong He , Jinlong Peng , Hao Yang , Mingmin Chi , Yabiao Wang

Precision livestock farming (PLF) increasingly relies on advanced object localization techniques to monitor livestock health and optimize resource management. This study investigates the generalization capabilities of YOLOv8 and YOLOv9…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Mautushi Das , Gonzalo Ferreira , C. P. James Chen

Wildlife camera trap images are being used extensively to investigate animal abundance, habitat associations, and behavior, which is complicated by the fact that experts must first classify the images manually. Artificial intelligence…

计算机视觉与模式识别 · 计算机科学 2023-08-03 Ludwig Bothmann , Lisa Wimmer , Omid Charrakh , Tobias Weber , Hendrik Edelhoff , Wibke Peters , Hien Nguyen , Caryl Benjamin , Annette Menzel

Rare object detection is a fundamental task in applied geospatial machine learning, however is often challenging due to large amounts of high-resolution satellite or aerial imagery and few or no labeled positive samples to start with. This…

Autonomous driving, particularly navigating complex and unanticipated scenarios, demands sophisticated reasoning and planning capabilities. While Multi-modal Large Language Models (MLLMs) offer a promising avenue for this, their use has…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Hidehisa Arai , Keita Miwa , Kento Sasaki , Yu Yamaguchi , Kohei Watanabe , Shunsuke Aoki , Issei Yamamoto

A multivariate time series refers to observations of two or more variables taken from a device or a system simultaneously over time. There is an increasing need to monitor multivariate time series and detect anomalies in real time to ensure…

机器学习 · 计算机科学 2023-05-29 Ming-Chang Lee , Jia-Chun Lin

We have extended our previous work to use the Murchison Widefield Array (MWA) as a non-coherent passive radar system in the FM frequency band, using terrestrial FM transmitters to illuminate objects in Low Earth Orbit LEO) and the MWA as…

天体物理仪器与方法 · 物理学 2020-12-16 Steve Prabu , Paul J Hancock , Zhang Xiang , Steven J Tingay

The evaluation of object detection models is usually performed by optimizing a single metric, e.g. mAP, on a fixed set of datasets, e.g. Microsoft COCO and Pascal VOC. Due to image retrieval and annotation costs, these datasets consist…

计算机视觉与模式识别 · 计算机科学 2022-12-01 Floriana Ciaglia , Francesco Saverio Zuppichini , Paul Guerrie , Mark McQuade , Jacob Solawetz

Reducing methane emissions is essential for mitigating global warming. To attribute methane emissions to their sources, a comprehensive dataset of methane source infrastructure is necessary. Recent advancements with deep learning on…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Bryan Zhu , Nicholas Lui , Jeremy Irvin , Jimmy Le , Sahil Tadwalkar , Chenghao Wang , Zutao Ouyang , Frankie Y. Liu , Andrew Y. Ng , Robert B. Jackson

Wildlife monitoring is crucial for studying biodiversity loss and climate change. Camera trap images provide a non-intrusive method for analyzing animal populations and identifying ecological patterns over time. However, manual analysis is…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Julian D. Santamaria , Claudia Isaza , Jhony H. Giraldo

We introduce Eagle 2.5, a family of frontier vision-language models (VLMs) for long-context multimodal learning. Our work addresses the challenges in long video comprehension and high-resolution image understanding, introducing a generalist…

Unmanned Aerial Vehicles (UAVs) have quickly become common in various airspaces, representing a wide range of applications from recreation flying to commercial photography and package delivery. With the increasing prevalence of UAVs, it…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Knut Peterson , Zaid Mayers , Azmain Yousuf , Priontu Chowdhury , Asher Zaczepinski , Solmaz Arezoomandan , Reihaneh Maarefdoust , David Han

While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hindered by a fundamental "temporal blindness". Existing architectures lack intrinsic…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Xiaohe Li , Jiahao Li , Kaixin Zhang , Yuqiang Fang , Leilei Lin , Hong Wang , Haohua Wu , Zide Fan

You Look Only Once (YOLO) models have been widely used for building real-time object detectors across various domains. With the increasing frequency of new YOLO versions being released, key questions arise. Are the newer versions always…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Tianyou Jiang , Yang Zhong

Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large…

Hornbills, an iconic species of Malaysia's biodiversity, face threats from habi-tat loss, poaching, and environmental changes, necessitating accurate and real-time population monitoring that is traditionally challenging and re-source…

声音 · 计算机科学 2025-04-17 Kong Ka Hing , Mehran Behjati

Precise localization and recognition of flowers are crucial for advancing automated agriculture, particularly in plant phenotyping, crop estimation, and yield monitoring. This paper benchmarks several YOLO architectures such as YOLOv5s,…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Safwat Nusrat , Prithwiraj Bhattacharjee

Camera traps have become integral tools in wildlife conservation, providing non-intrusive means to monitor and study wildlife in their natural habitats. The utilization of object detection algorithms to automate species identification from…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Aroj Subedi

Vision-Language-Action (VLA) models show strong generalization for robotic control, but finetuning them with reinforcement learning (RL) is constrained by the high cost and safety risks of real-world interaction. Training VLA models in…

机器人学 · 计算机科学 2026-03-24 Zhilong Zhang , Haoxiang Ren , Yihao Sun , Yifei Sheng , Haonan Wang , Haoxin Lin , Zhichao Wu , Pierre-Luc Bacon , Yang Yu