English
Related papers

Related papers: BlanketGen - A synthetic blanket occlusion augment…

200 papers

Inspired by generative paradigms in image and video, 3D shape generation has made notable progress, enabling the rapid synthesis of high-fidelity 3D assets from a single image. However, current methods still face challenges, including the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yangguang Li , Xianglong He , Zi-Xin Zou , Zexiang Liu , Wanli Ouyang , Ding Liang , Yan-Pei Cao

Data augmentation is a crucial technique in deep learning, particularly for tasks with limited dataset diversity, such as skeleton-based datasets. This paper proposes a comprehensive data augmentation framework that integrates geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Nada Aboudeshish , Dmitry Ignatov , Radu Timofte

Accurately annotated image datasets are essential components for studying animal behaviors from their poses. Compared to the number of species we know and may exist, the existing labeled pose datasets cover only a small portion of them,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Le Jiang , Shuangjun Liu , Xiangyu Bai , Sarah Ostadabbas

Despite advances in neural rendering, due to the scarcity of high-quality 3D datasets and the inherent limitations of multi-view diffusion models, view synthesis and 3D model generation are restricted to low resolutions with suboptimal…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Yihang Luo , Shangchen Zhou , Yushi Lan , Xingang Pan , Chen Change Loy

Despite recent success on 2D human pose estimation, 3D human pose estimation still remains an open problem. A key challenge is the ill-posed depth ambiguity nature. This paper presents a novel intermediate feature representation named…

Computer Vision and Pattern Recognition · Computer Science 2017-11-30 Qingfu Wan , Wei Zhang , Xiangyang Xue

We present DriveGen3D, a novel framework for generating high-quality and highly controllable dynamic 3D driving scenes that addresses critical limitations in existing methodologies. Current approaches to driving scene synthesis either…

Large Vision-Language Models (LVLMs) have shown promising capabilities in understanding and generating information by integrating both visual and textual data. However, current models are still prone to hallucinations, which degrade the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Robert Wijaya , Ngoc-Bao Nguyen , Ngai-Man Cheung

Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Feipeng Ma , Yizhou Zhou , Fengyun Rao , Yueyi Zhang , Xiaoyan Sun

We study how to synthesize novel views of human body from a single image. Though recent deep learning based methods work well for rigid objects, they often fail on objects with large articulation, like human bodies. The core step of…

Computer Vision and Pattern Recognition · Computer Science 2018-04-13 Hao Zhu , Hao Su , Peng Wang , Xun Cao , Ruigang Yang

Through automation, deep learning (DL) can enhance the analysis of transesophageal echocardiography (TEE) images. However, DL methods require large amounts of high-quality data to produce accurate results, which is difficult to satisfy.…

Image and Video Processing · Electrical Eng. & Systems 2024-10-10 Emmanuel Oladokun , Musa Abdulkareem , Jurica Šprem , Vicente Grau

Unsupervised learning of depth from indoor monocular videos is challenging as the artificial environment contains many textureless regions. Fortunately, the indoor scenes are full of specific structures, such as planes and lines, which…

Computer Vision and Pattern Recognition · Computer Science 2021-10-13 Hualie Jiang , Laiyan Ding , Junjie Hu , Rui Huang

Real-world 3D data may contain intricate details defined by salient surface gaps. Automated reconstruction of these open surfaces (e.g., non-watertight meshes) is a challenging problem for environment synthesis in mixed reality…

Computer Vision and Pattern Recognition · Computer Science 2023-01-20 Mohammad Samiul Arshad , William J. Beksi

Integrating human feedback in models can improve the performance of natural language processing (NLP) models. Feedback can be either explicit (e.g. ranking used in training language models) or implicit (e.g. using human cognitive signals in…

Human-Computer Interaction · Computer Science 2023-02-14 Varun Khurana , Yaman Kumar Singla , Nora Hollenstein , Rajesh Kumar , Balaji Krishnamurthy

Human behavioral monitoring during sleep is essential for various medical applications. Majority of the contactless human pose estimation algorithms are based on RGB modality, causing ineffectiveness in in-bed pose estimation due to…

Computer Vision and Pattern Recognition · Computer Science 2021-10-08 Mohamed Afham , Udith Haputhanthri , Jathurshan Pradeepkumar , Mithunjha Anandakumar , Ashwin De Silva , Chamira Edussooriya

The growing number of pretrained models in Machine Learning (ML) presents significant challenges for practitioners. Given a new dataset, they need to determine the most suitable deep learning (DL) pipeline, consisting of the pretrained…

Machine Learning · Computer Science 2025-06-17 Fabio Ferreira

The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from…

Computer Vision and Pattern Recognition · Computer Science 2024-05-06 Che-Jui Chang , Danrui Li , Deep Patel , Parth Goel , Honglu Zhou , Seonghyeon Moon , Samuel S. Sohn , Sejong Yoon , Vladimir Pavlovic , Mubbasir Kapadia

Datasets are crucial when training a deep neural network. When datasets are unrepresentative, trained models are prone to bias because they are unable to generalise to real world settings. This is particularly problematic for models trained…

Computer Vision and Pattern Recognition · Computer Science 2019-12-30 Mkhuseli Ngxande , Jules-Raymond Tapamo , Michael Burke

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Bowen Hao , Dongliang Zhou , Xiaojie Li , Xingyu Zhang , Liang Xie , Jianlong Wu , Erwei Yin

In this paper, a real-time method called PoP-Net is proposed to predict multi-person 3D poses from a depth image. PoP-Net learns to predict bottom-up part representations and top-down global poses in a single shot. Specifically, a new…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Yuliang Guo , Zhong Li , Zekun Li , Xiangyu Du , Shuxue Quan , Yi Xu

A key assumption of top-down human pose estimation approaches is their expectation of having a single person/instance present in the input bounding box. This often leads to failures in crowded scenes with occlusions. We propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Rawal Khirodkar , Visesh Chari , Amit Agrawal , Ambrish Tyagi