English
Related papers

Related papers: FixMyPose: Pose Correctional Captioning and Retrie…

200 papers

Image captioning bridges the gap between vision and language by automatically generating natural language descriptions for images. Traditional image captioning methods often overlook the preferences and characteristics of users.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Xuan Wang , Guanhong Wang , Wenhao Chai , Jiayu Zhou , Gaoang Wang

Most monocular and physics-based human pose tracking methods, while achieving state-of-the-art results, suffer from artifacts when the scene does not have a strictly flat ground plane or when the camera is moving. Moreover, these methods…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Ayce Idil Aytekin , Chuqiao Li , Diogo Luvizon , Rishabh Dabral , Martin Oswald , Marc Habermann , Christian Theobalt

Current vision-language models (VLMs) are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-27 Dewen Zhang , Tahir Hussain , Wangpeng An , Hayaru Shouno

Accurate human trajectory prediction is one of the most crucial tasks for autonomous driving, ensuring its safety. Yet, existing models often fail to fully leverage the visual cues that humans subconsciously communicate when navigating the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Yang Gao , Saeed Saadatnejad , Alexandre Alahi

Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Runyang Feng , Yixing Gao , Tze Ho Elden Tse , Xueqing Ma , Hyung Jin Chang

In this paper, we propose a novel approach to enhance the 3D body pose estimation of a person computed from videos captured from a single wearable camera. The key idea is to leverage high-level features linking first- and third-views in a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Ameya Dhamanaskar , Mariella Dimiccoli , Enric Corona , Albert Pumarola , Francesc Moreno-Noguer

We present Fitness tutor, an application for maintaining correct posture during workout exercises or doing yoga. Current work on fitness focuses on suggesting food supplements, accessing workouts, workout wearables does a great job in…

Computer Vision and Pattern Recognition · Computer Science 2021-09-06 Mahendran N

Full-body egocentric pose estimation from head and hand poses alone has become an active area of research to power articulate avatar representations on headset-based platforms. However, existing methods over-rely on the indoor…

Computer Vision and Pattern Recognition · Computer Science 2024-09-09 Jiaxi Jiang , Paul Streli , Manuel Meier , Christian Holz

This research presents the idea of activity fusion into existing Pose Estimation architectures to enhance their predictive ability. This is motivated by the rise in higher level concepts found in modern machine learning architectures, and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 David Poulton , Richard Klein

Dynamic multi-person mesh recovery has broad applications in sports broadcasting, virtual reality, and video games. However, current multi-view frameworks rely on a time-consuming camera calibration procedure. In this work, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Buzhen Huang , Jingyi Ju , Yuan Shu , Yangang Wang

In reading tasks drift can move fixations from one word to another or even another line, invalidating the eye tracking recording. Manual correction is time-consuming and subjective, while automated correction is fast yet limited in…

Human-Computer Interaction · Computer Science 2025-01-14 Naser Al Madi , Brett Torra , Yixin Li , Najam Tariq

Language is often used to describe physical interaction, yet most 3D human pose estimation methods overlook this rich source of information. We bridge this gap by leveraging large multimodal models (LMMs) as priors for reconstructing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Sanjay Subramanian , Evonne Ng , Lea Müller , Dan Klein , Shiry Ginosar , Trevor Darrell

Single-person human pose estimation facilitates markerless movement analysis in sports, as well as in clinical applications. Still, state-of-the-art models for human pose estimation generally do not meet the requirements of real-life…

Computer Vision and Pattern Recognition · Computer Science 2021-04-12 Daniel Groos , Heri Ramampiaro , Espen A. F. Ihlen

In this paper, we propose a new approach to learn multimodal multilingual embeddings for matching images and their relevant captions in two languages. We combine two existing objective functions to make images and captions close in a joint…

Computation and Language · Computer Science 2020-11-02 Alireza Mohammadshahi , Remi Lebret , Karl Aberer

Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention. However, existing datasets for monocular pose estimation do not adequately capture the challenging and dynamic nature of sports movements.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Christian Keilstrup Ingwersen , Christian Mikkelstrup , Janus Nørtoft Jensen , Morten Rieger Hannemose , Anders Bjorholm Dahl

Human pose estimation and tracking are fundamental tasks for understanding human behaviors in videos. Existing top-down framework-based methods usually perform three-stage tasks: human detection, pose estimation and tracking. Although…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zehua Fu , Wenhang Zuo , Zhenghui Hu , Qingjie Liu , Yunhong Wang

Visually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision. However, modern audio-visual datasets contain biases that undermine the real-world performance of models trained…

Computation and Language · Computer Science 2021-10-15 Ian Palmer , Andrew Rouditchenko , Andrei Barbu , Boris Katz , James Glass

Human pose estimation from single images is a challenging problem in computer vision that requires large amounts of labeled training data to be solved accurately. Unfortunately, for many human activities (\eg outdoor sports) such training…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Bastian Wandt , Marco Rudolph , Petrissa Zell , Helge Rhodin , Bodo Rosenhahn

Gaze reflects how humans process visual scenes and is therefore increasingly used in computer vision systems. Previous works demonstrated the potential of gaze for object-centric tasks, such as object localization and recognition, but it…

Computer Vision and Pattern Recognition · Computer Science 2016-08-19 Yusuke Sugano , Andreas Bulling

This paper presents Text2Traj2Text, a novel learning-by-synthesis framework for captioning possible contexts behind shopper's trajectory data in retail stores. Our work will impact various retail applications that need better customer…

Computation and Language · Computer Science 2024-09-20 Hikaru Asano , Ryo Yonetani , Taiki Sekii , Hiroki Ouchi
‹ Prev 1 4 5 6 7 8 10 Next ›