English
Related papers

Related papers: AnimalFormer: Multimodal Vision Framework for Beha…

200 papers

Large AI models have been widely adopted in wireless communications for channel modeling, beamforming, and resource optimization. However, most existing efforts remain limited to single-modality inputs and channel-specific objec- tives,…

Machine Learning · Computer Science 2025-11-18 Zhizhen Li , Xuanhao Luo , Xueren Ge , Longyu Zhou , Xingqin Lin , Yuchen Liu

Referring expression grounding is an important and challenging task in computer vision. To avoid the laborious annotation in conventional referring grounding, unpaired referring grounding is introduced, where the training data only contains…

Computer Vision and Pattern Recognition · Computer Science 2022-06-07 Hengcan Shi , Munawar Hayat , Jianfei Cai

The advent of high-resolution multispectral/hyperspectral sensors, LiDAR DSM (Digital Surface Model) information and many others has provided us with an unprecedented wealth of data for Earth Observation. Multimodal AI seeks to exploit…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Nhi Kieu , Kien Nguyen , Sridha Sridharan , Clinton Fookes

Visual Odometry (VO) estimation is an important source of information for vehicle state estimation and autonomous driving. Recently, deep learning based approaches have begun to appear in the literature. However, in the context of driving,…

Computer Vision and Pattern Recognition · Computer Science 2021-12-28 Nimet Kaygusuz , Oscar Mendez , Richard Bowden

Achieving robust vision-based humanoid locomotion remains challenging due to two fundamental issues: the sim-to-real gap introduces significant perception noise that degrades performance on fine-grained tasks, and training a unified policy…

Animal tracking and pose estimation systems, such as STEP (Simultaneous Tracking and Pose Estimation) and ViTPose, experience substantial performance drops when processing images and videos with cage structures and systematic occlusions. We…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Sayak Dutta , Harish Katti , Shashikant Verma , Shanmuganathan Raman

We study active object tracking, where a tracker takes visual observations (i.e., frame sequences) as input and produces the corresponding camera control signals as output (e.g., move forward, turn left, etc.). Conventional methods tackle…

Computer Vision and Pattern Recognition · Computer Science 2019-02-14 Wenhan Luo , Peng Sun , Fangwei Zhong , Wei Liu , Tong Zhang , Yizhou Wang

The optimisation of crop harvesting processes for commonly cultivated crops is of great importance in the aim of agricultural industrialisation. Nowadays, the utilisation of machine vision has enabled the automated identification of crops,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Hongyu Zhao , Zezhi Tang , Zhenhong Li , Yi Dong , Yuancheng Si , Mingyang Lu , George Panoutsos

Iris segmentation is the initial step to identify biometric of animals to establish a traceability system of livestock. In this study, we propose a novel deep learning framework for pixel-wise segmentation with minimum use of annotation…

Image and Video Processing · Electrical Eng. & Systems 2022-12-23 Heemoon Yoon , Mira Park , Sang-Hee Lee

The model-based estimation of 3D animal pose and shape from images enables computational modeling of animal behavior. Training models for this purpose requires large amounts of labeled image data with precise pose and shape annotations.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Tomasz Niewiadomski , Anastasios Yiannakidis , Hanz Cuevas-Velasquez , Soubhik Sanyal , Michael J. Black , Silvia Zuffi , Peter Kulits

We introduce SkelFormer, a novel markerless motion capture pipeline for multi-view human pose and shape estimation. Our method first uses off-the-shelf 2D keypoint estimators, pre-trained on large-scale in-the-wild data, to obtain 3D joint…

Computer Vision and Pattern Recognition · Computer Science 2024-04-22 Vandad Davoodnia , Saeed Ghorbani , Alexandre Messier , Ali Etemad

Language-instructed robot manipulation has garnered significant interest due to the potential of learning from collected data. While the challenges in high-level perception and planning are continually addressed along the progress of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Shanshan Guo , Xiwen Liang , Junfan Lin , Yuzheng Zhuang , Liang Lin , Xiaodan Liang

In precision sports such as archery, athletes' performance depends on both biomechanical stability and psychological resilience. Traditional motion analysis systems are often expensive and intrusive, limiting their use in natural training…

Machine Learning · Computer Science 2025-11-19 Xianghe Liu , Jiajia Liu , Chuxian Xu , Minghan Wang , Hongbo Peng , Tao Sun , Jiaqi Xu

Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Daixun Li , Weiying Xie , Mingxiang Cao , Yunke Wang , Yusi Zhang , Leyuan Fang , Yunsong Li , Chang Xu

Self-supervised learning (SSL) leverages vast unannotated medical datasets, yet steep technical barriers limit adoption by clinical researchers. We introduce Vision Foundry, a code-free, HIPAA-compliant platform that democratizes…

Markerless methods for animal posture tracking have been rapidly developing recently, but frameworks and benchmarks for tracking large animal groups in 3D are still lacking. To overcome this gap in the literature, we present 3D-MuPPET, a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 Urs Waldmann , Alex Hoi Hang Chan , Hemal Naik , Máté Nagy , Iain D. Couzin , Oliver Deussen , Bastian Goldluecke , Fumihiro Kano

We introduce a novel method for human shape and pose recovery that can fully leverage multiple static views. We target fixed-multiview people monitoring, including elderly care and safety monitoring, in which calibrated cameras can be…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yuto Matsubara , Ko Nishino

We introduce a unified, end-to-end framework that seamlessly integrates object detection and pose estimation with a versatile onboarding process. Our pipeline begins with an onboarding stage that generates object representations from either…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Artem Moroz , Vít Zeman , Martin Mikšík , Elizaveta Isianova , Miroslav David , Pavel Burget , Varun Burde

Addressing the current lack of a standardized habitat classification system for cultivated land ecosystems, incomplete coverage of the habitat types, and the inability of existing models to effectively integrate semantic and texture…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Kesong Zheng , Zhi Song , Peizhou Li , Shuyi Yao , Zhenxing Bian

Visual grounding aims to identify objects or regions in a scene based on natural language descriptions, essential for spatially aware perception in autonomous driving. However, existing visual grounding tasks typically depend on bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Zhan Shi , Song Wang , Junbo Chen , Jianke Zhu