English
Related papers

Related papers: PointOdyssey: A Large-Scale Synthetic Dataset for …

200 papers

Data augmentation methods such as Copy-Paste have been studied as effective ways to expand training datasets while incurring minimal costs. While such methods have been extensively implemented for image level tasks, we found no scalable…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Sahir Shrestha , Weihao Li , Gao Zhu , Nick Barnes

Data-efficient training of robust robot policies is the key to unlocking automation in a wide array of novel tasks. Current systems require large volumes of demonstrations to achieve robustness, which is impractical in many applications.…

Robotics · Computer Science 2026-03-10 Adam Hung , Bardienus Pieter Duisterhof , Jeffrey Ichnowski

Recent advancements in diffusion models have significantly improved the realism and generalizability of character-driven animation, enabling the synthesis of high-quality motion from just a single RGB image and a set of driving poses.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Alireza Javanmardi , Pragati Jaiswal , Tewodros Amberbir Habtegebrial , Christen Millerdurai , Shaoxiang Wang , Alain Pagani , Didier Stricker

This paper addresses the problem of 3D human pose estimation in the wild. A significant challenge is the lack of training data, i.e., 2D images of humans annotated with 3D poses. Such data is necessary to train state-of-the-art CNN…

Computer Vision and Pattern Recognition · Computer Science 2018-02-13 Grégory Rogez , Cordelia Schmid

Controllable video synthesis is a central challenge in computer vision, yet current models struggle with fine grained control beyond textual prompts, particularly for cinematic attributes like camera trajectory and genre. Existing datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Zahra Dehghanian , Morteza Abolghasemi , Hamid Beigy , Hamid R. Rabiee

We present an overview and evaluation of a new, systematic approach for generation of highly realistic, annotated synthetic data for training of deep neural networks in computer vision tasks. The main contribution is a procedural world…

Computer Vision and Pattern Recognition · Computer Science 2017-10-19 Apostolia Tsirikoglou , Joel Kronander , Magnus Wrenninge , Jonas Unger

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Shuolin Xu , Siming Zheng , Ziyi Wang , HC Yu , Jinwei Chen , Huaqi Zhang , Daquan Zhou , Tong-Yee Lee , Bo Li , Peng-Tao Jiang

Human video synthesis aims to create lifelike characters in various environments, with wide applications in VR, storytelling, and content creation. While 2D diffusion-based methods have made significant progress, they struggle to generalize…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Liyuan Cui , Xiaogang Xu , Wenqi Dong , Zesong Yang , Hujun Bao , Zhaopeng Cui

Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention. However, existing datasets for monocular pose estimation do not adequately capture the challenging and dynamic nature of sports movements.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Christian Keilstrup Ingwersen , Christian Mikkelstrup , Janus Nørtoft Jensen , Morten Rieger Hannemose , Anders Bjorholm Dahl

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Chin-Yang Lin , Cheng Sun , Fu-En Yang , Min-Hung Chen , Yen-Yu Lin , Yu-Lun Liu

Ultrasound (US) is widely used for its advantages of real-time imaging, radiation-free and portability. In clinical practice, analysis and diagnosis often rely on US sequences rather than a single image to obtain dynamic anatomical…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Jiamin Liang , Xin Yang , Yuhao Huang , Kai Liu , Xinrui Zhou , Xindi Hu , Zehui Lin , Huanjia Luo , Yuanji Zhang , Yi Xiong , Dong Ni

Despite increasingly realistic image quality, recent 3D image generative models often operate on 3D volumes of fixed extent with limited camera motions. We investigate the task of unconditionally synthesizing unbounded nature scenes,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Lucy Chai , Richard Tucker , Zhengqi Li , Phillip Isola , Noah Snavely

In this paper, we present LaSOT, a high-quality benchmark for Large-scale Single Object Tracking. LaSOT consists of 1,400 sequences with more than 3.5M frames in total. Each frame in these sequences is carefully and manually annotated with…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Heng Fan , Liting Lin , Fan Yang , Peng Chu , Ge Deng , Sijia Yu , Hexin Bai , Yong Xu , Chunyuan Liao , Haibin Ling

Structured 3D representations such as keypoints and meshes offer compact, expressive descriptions of deformable objects, jointly capturing geometric and topological information useful for downstream tasks such as dynamics modeling and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Yeheng Zong , Yizhou Chen , Alexander Bowler , Chia-Tung Yang , Ram Vasudevan

Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and foreground annotations, which must be composited to background…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Yongtao Ge , Kangyang Xie , Guangkai Xu , Mingyu Liu , Li Ke , Longtao Huang , Hui Xue , Hao Chen , Chunhua Shen

This work presents a novel video dataset recorded from overlapping highway traffic cameras along an urban interstate, enabling multi-camera 3D object tracking in a traffic monitoring context. Data is released from 3 scenes containing video…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Derek Gloudemans , Yanbing Wang , Gracie Gumm , William Barbour , Daniel B. Work

We introduce the dynamic grasp synthesis task: given an object with a known 6D pose and a grasp reference, our goal is to generate motions that move the object to a target 6D pose. This is challenging, because it requires reasoning about…

Computer Vision and Pattern Recognition · Computer Science 2022-04-18 Sammy Christen , Muhammed Kocabas , Emre Aksan , Jemin Hwangbo , Jie Song , Otmar Hilliges

Extracting physical dynamical system parameters from recorded observations is key in natural science. Current methods for automatic parameter estimation from video train supervised deep networks on large datasets. Such datasets require…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Alejandro Castañeda Garcia , Jan van Gemert , Daan Brinks , Nergis Tömen

Video data is more cost-effective than motion capture data for learning 3D character motion controllers, yet synthesizing realistic and diverse behaviors directly from videos remains challenging. Previous approaches typically rely on…

Graphics · Computer Science 2025-12-10 Jianan Li , Xiao Chen , Tao Huang , Tien-Tsin Wong

This paper addresses the challenges of data scarcity and high acquisition costs in training robust object detection models for complex industrial environments, such as offshore oil platforms. Data collection in these hazardous settings…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Pedro Antonio Rabelo Saraiva , Enzo Ferreira de Souza , Joao Manoel Herrera Pinheiro , Thiago H. Segreto , Ricardo V. Godoy , Marcelo Becker