中文
相关论文

相关论文: AP-CAP: Advancing High-Quality Data Synthesis for …

200 篇论文

We propose the first metric learning system for the recognition of great ape behavioural actions. Our proposed triple stream embedding architecture works on camera trap videos taken directly in the wild and demonstrates that the utilisation…

计算机视觉与模式识别 · 计算机科学 2023-01-09 Otto Brookes , Majid Mirmehdi , Hjalmar Kühl , Tilo Burghardt

Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While several datasets exist for natural images and high-resolution optical remote sensing…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Lucrezia Tosato , Gianluca Lombardi , Ronny Hansch

Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts which are costly to acquire. Previous works leverage the…

计算机视觉与模式识别 · 计算机科学 2023-12-15 Zhiyue Liu , Jinyuan Liu , Fanrong Ma

In this paper, we are interested in pose estimation of animals. Animals usually exhibit a wide range of variations on poses and there is no available animal pose dataset for training and testing. To address this problem, we build an animal…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Jinkun Cao , Hongyang Tang , Hao-Shu Fang , Xiaoyong Shen , Cewu Lu , Yu-Wing Tai

Person re-identification (person Re-Id) aims to retrieve the pedestrian images of a same person that captured by disjoint and non-overlapping cameras. Lots of researchers recently focuse on this hot issue and propose deep learning based…

计算机视觉与模式识别 · 计算机科学 2019-06-06 Chengyuan Zhang , Lei Zhu , Shichao Zhang

Multi-view pose estimation is essential for quantifying animal behavior in scientific research, yet current methods struggle to achieve accurate tracking with limited labeled data and suffer from poor uncertainty estimates. We address these…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Lenny Aharon , Keemin Lee , Karan Sikka , Selmaan Chettih , Cole Hurwitz , Liam Paninski , Matthew R Whiteway

In this letter, we present a novel markerless 3D human motion capture (MoCap) system for unstructured, outdoor environments that uses a team of autonomous unmanned aerial vehicles (UAVs) with on-board RGB cameras and computation. Existing…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Nitin Saini , Elia Bonetto , Eric Price , Aamir Ahmad , Michael J. Black

Pose diversity is an inherent representative characteristic of 2D images. Due to the 3D to 2D projection mechanism, there is evident content discrepancy among distinct pose images. This is the main obstacle bothering pose transformation…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Yuelong Li , Tengfei Xiao , Lei Geng , Jianming Wang

Numerous fields, such as ecology, biology, and neuroscience, use animal recordings to track and measure animal behaviour. Over time, a significant volume of such data has been produced, but some computer vision techniques cannot explore it…

计算机视觉与模式识别 · 计算机科学 2023-07-26 Jose Sosa , Sharn Perry , Jane Alty , David Hogg

Animal pose estimation has recently come into the limelight due to its application in biology, zoology, and aquaculture. Deep learning methods have effectively been applied to human pose estimation. However, the major bottleneck to the…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Samayan Bhattacharya , Sk Shahnawaz

Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware and markers limits scalability and real-world deployment. Advancing reliable markerless…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Yeeun Park , Miqdad Naduthodi , Suryansh Kumar

Acquiring labeled datasets for 3D human mesh estimation is challenging due to depth ambiguities and the inherent difficulty of annotating 3D geometry from monocular images. Existing datasets are either real, with manually annotated 3D…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Lorenza Prospero , Orest Kupyn , Ostap Viniavskyi , João F. Henriques , Christian Rupprecht

The effectiveness of Contrastive Language-Image Pre-training (CLIP) models critically depends on the semantic diversity and quality of their training data. However, while existing synthetic data generation methods primarily focus on…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Yuanxiang Huangfu , Chaochao Wang , Weilei Wang

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains…

计算机视觉与模式识别 · 计算机科学 2019-01-24 Jiefeng Li , Can Wang , Hao Zhu , Yihuan Mao , Hao-Shu Fang , Cewu Lu

This paper addresses the problem of cross-dataset generalization of 3D human pose estimation models. Testing a pre-trained 3D pose estimator on a new dataset results in a major performance drop. Previous methods have mainly addressed this…

计算机视觉与模式识别 · 计算机科学 2022-03-17 Mohsen Gholami , Bastian Wandt , Helge Rhodin , Rabab Ward , Z. Jane Wang

Existing research on unconstrained in-the-wild head pose estimation suffers from the flaws of its datasets, which consist of either numerous samples by non-realistic synthesis or constrained collection, or small-scale natural images yet…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Huayi Zhou , Fei Jiang , Jin Yuan , Yong Rui , Hongtao Lu , Kui Jia

Automated capture of animal pose is transforming how we study neuroscience and social behavior. Movements carry important social cues, but current methods are not able to robustly estimate pose and shape of animals, particularly for social…

计算机视觉与模式识别 · 计算机科学 2021-01-13 Marc Badger , Yufu Wang , Adarsh Modh , Ammon Perkes , Nikos Kolotouros , Bernd G. Pfrommer , Marc F. Schmidt , Kostas Daniilidis

Accurate 6D pose estimation of complex objects in 3D environments is essential for effective robotic manipulation. Yet, existing benchmarks fall short in evaluating 6D pose estimation methods under realistic industrial conditions, as most…

In this paper, we introduce Context-Aware Priority Sampling (CAPS), a novel method designed to enhance data efficiency in learning-based autonomous driving systems. CAPS addresses the challenge of imbalanced datasets in imitation learning…

6D pose estimation refers to object recognition and estimation of 3D rotation and 3D translation. The key technology for estimating 6D pose is to estimate pose by extracting enough features to find pose in any environment. Previous methods…

计算机视觉与模式识别 · 计算机科学 2020-08-13 Myoungha Song , Jeongho Lee , Donghwan Kim