中文
相关论文

相关论文: AnimalFormer: Multimodal Vision Framework for Beha…

200 篇论文

Recently significant progress has been made in human action recognition and behavior prediction using deep learning techniques, leading to improved vision-based semantic understanding. However, there is still a lack of high-quality motion…

计算机视觉与模式识别 · 计算机科学 2023-07-24 Xiaofeng Liu , Jiaxin Gao , Yaohua Liu , Risheng Liu , Nenggan Zheng

Self-supervised learning has emerged as a powerful paradigm for training deep neural networks, particularly in medical imaging where labeled data is scarce. While current approaches typically rely on synthetic augmentations of single…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Andre Dourson , Kylie Taylor , Xiaoli Qiao , Michael Fitzke

Assessing human skill levels in complex activities is a challenging problem with applications in sports, rehabilitation, and training. In this work, we present SkillFormer, a parameter-efficient architecture for unified multi-view…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Edoardo Bianchi , Antonio Liotta

This study presents a vision-guided robotic control system for automated fruit tree pruning applications. Traditional pruning practices are labor-intensive and limit agricultural efficiency and scalability, highlighting the need for…

机器人学 · 计算机科学 2025-06-09 Dawood Ahmed , Basit Muhammad Imran , Martin Churuvija , Manoj Karkee

While Visual Servoing is deeply studied to perform simple maneuvers, the literature does not commonly address complex cases where the target is far out of the camera's field of view (FOV) during the maneuver. For this reason, in this paper,…

Traditional monitoring of bearded dragon (Pogona Viticeps) behaviour is time-consuming and prone to errors. This project introduces an automated system for real-time video analysis, using You Only Look Once (YOLO) object detection models to…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Arsen Yermukan , Pedro Machado , Feliciano Domingos , Isibor Kennedy Ihianle , Jordan J. Bird , Stefano S. K. Kaburu , Samantha J. Ward

We tackle the challenges of synthesizing versatile, physically simulated human motions for full-body object manipulation. Unlike prior methods that are focused on detailed motion tracking, trajectory following, or teleoperation, our…

机器人学 · 计算机科学 2025-12-12 Chen Tessler , Yifeng Jiang , Erwin Coumans , Zhengyi Luo , Gal Chechik , Xue Bin Peng

We present WeedRepFormer, a lightweight multi-task Vision Transformer designed for simultaneous waterhemp segmentation and gender classification. Existing agricultural models often struggle to balance the fine-grained feature extraction…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Toqi Tahamid Sarker , Taminul Islam , Khaled R. Ahmed , Cristiana Bernardi Rankrape , Kaitlin E. Creager , Karla Gage

Dense image segmentation tasks e.g., semantic, panoptic) are useful for image editing, but existing methods can hardly generalize well in an in-the-wild setting where there are unrestricted image domains, classes, and image resolution and…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Lu Qi , Jason Kuen , Weidong Guo , Tiancheng Shen , Jiuxiang Gu , Jiaya Jia , Zhe Lin , Ming-Hsuan Yang

Predicting pedestrian behavior is a crucial task for intelligent driving systems. Accurate predictions require a deep understanding of various contextual elements that potentially impact the way pedestrians behave. To address this…

计算机视觉与模式识别 · 计算机科学 2022-10-17 Amir Rasouli , Iuliia Kotseruba

Vision foundation models (VFMs) and Bird's Eye View (BEV) representation have advanced visual perception substantially, yet their internal spatial representations assume the rectilinear geometry of pinhole cameras. Fisheye cameras, widely…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Rahul Ahuja , Mudit Jain , Bala Murali Manoghar Sai Sudhakar , Venkatraman Narayanan , Pratik Likhar , Varun Ravi Kumar , Senthil Yogamani

Estimating camera pose in dynamic environments is a critical challenge, as most visual SLAM and SfM methods assume static scenes. While recent dynamic-aware methods exist, they are often not unified: semantic-based approaches are brittle,…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Jianhao Zheng , Liyuan Zhu , Zihan Zhu , Iro Armeni

Acquiring aligned visuo-tactile datasets is slow and costly, requiring specialised hardware and large-scale data collection. Synthetic generation is promising, but prior methods are typically single-modality, limiting cross-modal learning.…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Sirine Bhouri , Lan Wei , Jian-Qing Zheng , Dandan Zhang

Using UAVs for wildlife observation and motion capture offers manifold advantages for studying animals in the wild, especially grazing herds in open terrain. The aerial perspective allows observation at a scale and depth that is not…

机器人学 · 计算机科学 2024-05-27 Eric Price , Aamir Ahmad

Visual transformers have driven major progress in remote sensing image analysis, particularly in object detection and segmentation. Recent vision-language and multimodal models further extend these capabilities by incorporating auxiliary…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Yu Li , Guilherme N. DeSouza , Praveen Rao , Chi-Ren Shyu

Humans are excellent at understanding language and vision to accomplish a wide range of tasks. In contrast, creating general instruction-following embodied agents remains a difficult challenge. Prior work that uses pure language-only models…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Hao Liu , Lisa Lee , Kimin Lee , Pieter Abbeel

This study presents an innovative computer vision framework designed to analyze human movements in industrial settings, aiming to enhance biomechanical analysis by integrating seamlessly with existing software. Through a combination of…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Jesudara Omidokun , Darlington Egeonu , Bochen Jia , Liang Yang

Facial Action Unit (AU) detection in in-the-wild environments remains a formidable challenge due to severe spatial-temporal heterogeneity, unconstrained poses, and complex audio-visual dependencies. While recent multimodal approaches have…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Jun Yu , Yunxiang Zhang , Naixiang Zheng , Lingsi Zhu , Guoyuan Wang

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geographic coverage, or…

Advancements in machine vision that enable detailed inferences to be made from images have the potential to transform many sectors including agriculture. Precision agriculture, where data analysis enables interventions to be precisely…

计算机视觉与模式识别 · 计算机科学 2023-10-11 Madeleine Darbyshire , Elizabeth Sklar , Simon Parsons
‹ 上一页 1 8 9 10 下一页 ›