English
Related papers

Related papers: AnimalFormer: Multimodal Vision Framework for Beha…

200 papers

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across multimodal tasks such as visual perception and reasoning, leading to good performance on various multimodal evaluation benchmarks. However, these…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yue Yang , Shuibai Zhang , Wenqi Shao , Kaipeng Zhang , Yi Bin , Yu Wang , Ping Luo

Although no specific domain knowledge is considered in the design, plain vision transformers have shown excellent performance in visual recognition tasks. However, little effort has been made to reveal the potential of such simple…

Computer Vision and Pattern Recognition · Computer Science 2022-10-14 Yufei Xu , Jing Zhang , Qiming Zhang , Dacheng Tao

Human pose estimation, a vital task in computer vision, involves detecting and localising human joints in images and videos. While single-frame pose estimation has seen significant progress, it often fails to capture the temporal dynamics…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Cesare Davide Pace , Alessandro Marco De Nunzio , Claudio De Stefano , Francesco Fontanella , Mario Molinara

In the era of foundation models, achieving a unified understanding of different dynamic objects through a single network has the potential to empower stronger spatial intelligence. Moreover, accurate estimation of animal pose and shape…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Liang An , Jin Lyu , Li Lin , Pujin Cheng , Yebin Liu , Xiaoying Tang

Using drones to track multiple individuals simultaneously in their natural environment is a powerful approach for better understanding group primate behavior. Previous studies have demonstrated that it is possible to automate the…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Isla Duporge , Maksim Kholiavchenko , Roi Harel , Scott Wolf , Dan Rubenstein , Meg Crofoot , Tanya Berger-Wolf , Stephen Lee , Julie Barreau , Jenna Kline , Michelle Ramirez , Charles Stewart

Understanding animal behavior from video is essential for neuroscience research. Modern laboratories typically collect two complementary data streams: skeletal keypoints from pose estimation tools and raw video recordings. Keypoint-based…

Quantitative Methods · Quantitative Biology 2026-03-10 Weihan Li , Jingyang Ke , Yule Wang , Chengrui Li , Anqi Wu

Animal welfare has become a critical issue in contemporary society, emphasizing our ethical responsibilities toward animals, particularly within livestock farming. The advent of Artificial Intelligence (AI) technologies, specifically…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Voncarlos M. Araújo , Ines Rili , Thomas Gisiger , Sebastien Gambs , Elsa Vasseur , Marjorie Cellier , Abdoulaye Baniré Diallo

We introduce VidTFS, a Training-free, open-vocabulary video goal and action inference framework that combines the frozen vision foundational model (VFM) and large language model (LLM) with a novel dynamic Frame Selection module. Our…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Ee Yeo Keat , Zhang Hao , Alexander Matyasko , Basura Fernando

Crop monitoring is essential for precision agriculture, but current systems lack high-level reasoning. We introduce a novel, modular framework that uses a Visual Language Model (VLM) to guide robotic task planning, interleaving input…

Robotics · Computer Science 2026-01-21 Jose Cuaran , Kendall Koe , Aditya Potnis , Naveen Kumar Uppalapati , Girish Chowdhary

Autonomous vehicles (AVs) are poised to redefine transportation by enhancing road safety, minimizing human error, and optimizing traffic efficiency. The success of AVs depends on their ability to interpret complex, dynamic environments…

Multimedia · Computer Science 2025-07-11 Abolfazl Zarghani , Amirhossein Ebrahimi , Amir Malekesfandiari

Computer vision for animals holds great promise for wildlife research but often depends on large-scale data, while existing collection methods rely on controlled capture setups. Recent data-driven approaches show the potential of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Brian Nlong Zhao , Jiajun Wu , Shangzhe Wu

State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous visual tokens, many of which are irrelevant to the downstream task. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xinyu Huang , Yuhao Dong , Weiwei Tian , Bo Li , Rui Feng , Ziwei Liu

Meta-learning hyperparameter optimization (HPO) algorithms from prior experiments is a promising approach to improve optimization efficiency over objective functions from a similar distribution. However, existing methods are restricted to…

Understanding animals' behaviors is significant for a wide range of applications. However, existing animal behavior datasets have limitations in multiple aspects, including limited numbers of animal classes, data samples and provided tasks,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-06 Xun Long Ng , Kian Eng Ong , Qichen Zheng , Yun Ni , Si Yong Yeo , Jun Liu

The simplicity of the visual servoing approach makes it an attractive option for tasks dealing with vision-based control of robots in many real-world applications. However, attaining precise alignment for unseen environments pose a…

Grounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC). Its effectiveness has led to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 Xiangyu Zhao , Yicheng Chen , Shilin Xu , Xiangtai Li , Xinjiang Wang , Yining Li , Haian Huang

Perceptive locomotion for legged robots requires anticipating and adapting to complex, dynamic environments. Model Predictive Control (MPC) serves as a strong baseline, providing interpretable motion planning with constraint enforcement,…

Robotics · Computer Science 2026-03-17 Aditya Shirwatkar , Satyam Gupta , Shishir Kolathaya

Predicting human motion is critical for assistive robots and AR/VR applications, where the interaction with humans needs to be safe and comfortable. Meanwhile, an accurate prediction depends on understanding both the scene context and human…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Yang Zheng , Yanchao Yang , Kaichun Mo , Jiaman Li , Tao Yu , Yebin Liu , C. Karen Liu , Leonidas J. Guibas

This paper presents a novel reinforcement learning (RL)-based planning scheme for optimized robotic management of biotic stresses in precision agriculture. The framework employs a hierarchical decision-making structure with conditional…

We present a cross robot visuomotor learning framework that integrates diffusion policy based control with 3D semantic scene representations from D3Fields to enable category level generalization in manipulation. Its modular design supports…