English
Related papers

Related papers: MGAug: Multimodal Geometric Augmentation in Latent…

200 papers

The ability to jointly learn from multiple modalities, such as text, audio, and visual data, is a defining feature of intelligent systems. While there have been promising advances in designing neural networks to harness multimodal data, the…

Machine Learning · Computer Science 2023-04-25 Zichang Liu , Zhiqiang Tang , Xingjian Shi , Aston Zhang , Mu Li , Anshumali Shrivastava , Andrew Gordon Wilson

Deep generative models make visual content creation more accessible to novice users by automating the synthesis of diverse, realistic content based on a collected dataset. However, the current machine learning approaches miss a key element…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Sheng-Yu Wang , David Bau , Jun-Yan Zhu

Whole slide image (WSI) analysis in digital pathology presents unique challenges due to the gigapixel resolution of WSIs and the scarcity of dense supervision signals. While Multiple Instance Learning (MIL) is a natural fit for slide-level…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Sofiène Boutaj , Marin Scalbert , Pierre Marza , Florent Couzinie-Devy , Maria Vakalopoulou , Stergios Christodoulidis

In the past five years, research has shifted from traditional Machine Learning (ML) and Deep Learning (DL) approaches to leveraging Large Language Models (LLMs) , including multimodality, for data augmentation to enhance generalization, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Ranjan Sapkota , Shaina Raza , Maged Shoman , Achyut Paudel , Manoj Karkee

Generative adversarial networks (GANs) are unsupervised Deep Learning approach in the computer vision community which has gained significant attention from the last few years in identifying the internal structure of multimodal medical…

Image and Video Processing · Electrical Eng. & Systems 2020-05-22 Nripendra Kumar Singh , Khalid Raza

Data augmentation is a crucial technique for training robust deep learning models for human motion, where annotated datasets are often scarce. However, generic augmentation methods often ignore the underlying geometric and kinematic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Bikram De , Habib Irani , Vangelis Metsis

One of the growing trends in machine learning is the use of data generation techniques, since the performance of machine learning models is dependent on the quantity of the training dataset. However, in many real-world applications,…

Artificial Intelligence · Computer Science 2025-04-25 Yasaman Haghbin , Hadi Moradi , Reshad Hosseini

Reconstructing dynamic objects from monocular videos is a severely underconstrained and challenging problem, and recent work has approached it in various directions. However, owing to the ill-posed nature of this problem, there has been no…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Devikalyan Das , Christopher Wewer , Raza Yunus , Eddy Ilg , Jan Eric Lenssen

The goal of multimodal alignment is to learn a single latent space that is shared between multimodal inputs. The most powerful models in this space have been trained using massive datasets of paired inputs and large-scale computational…

Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a…

Computer Vision and Pattern Recognition · Computer Science 2022-09-08 Tomer Amit , Tal Shaharbany , Eliya Nachmani , Lior Wolf

Bimanual manipulation requires policies that can reason about 3D geometry, anticipate how it evolves under action, and generate smooth, coordinated motions. However, existing methods typically rely on 2D features with limited spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Chongyang Xu , Haipeng Li , Shen Cheng , Jingyu Hu , Haoqiang Fan , Ziliang Feng , Shuaicheng Liu

Image denoising is a fundamental problem in computational photography, where achieving high perception with low distortion is highly demanding. Current methods either struggle with perceptual quality or suffer from significant distortion.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Tong Li , Hansen Feng , Lizhi Wang , Zhiwei Xiong , Hua Huang

Diffusion models have shown remarkable flexibility for solving inverse problems without task-specific retraining. However, existing approaches such as Manifold Preserving Guided Diffusion (MPGD) apply only a single gradient update per…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Aditya Chakravarty

Dynamic garment reconstruction from monocular video is an important yet challenging task due to the complex dynamics and unconstrained nature of the garments. Recent advancements in neural rendering have enabled high-quality geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Soham Dasgupta , Shanthika Naik , Preet Savalia , Sujay Kumar Ingle , Avinash Sharma

Finding a realistic deformation that transforms one image into another, in case large deformations are required, is considered a key challenge in medical image analysis. Having a proper image registration approach to achieve this could…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Georgios Andreadis , Peter A. N. Bosman , Tanja Alderliesten

Convolutional neural networks have been widely applied to medical image segmentation and have achieved considerable performance. However, the performance may be significantly affected by the domain gap between training data (source domain)…

Image and Video Processing · Electrical Eng. & Systems 2022-07-28 Junyan Lyu , Yiqi Zhang , Yijin Huang , Li Lin , Pujin Cheng , Xiaoying Tang

Data augmentation has recently emerged as an essential component of modern training recipes for visual recognition tasks. However, data augmentation for video recognition has been rarely explored despite its effectiveness. Few existing…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Taeoh Kim , Jinhyung Kim , Minho Shim , Sangdoo Yun , Myunggu Kang , Dongyoon Wee , Sangyoun Lee

In the field of medical image, deep convolutional neural networks(ConvNets) have achieved great success in the classification, segmentation, and registration tasks thanks to their unparalleled capacity to learn image features. However,…

Image and Video Processing · Electrical Eng. & Systems 2022-11-15 Xin Gao

Data augmentation is a crucial technique in deep learning, particularly for tasks with limited dataset diversity, such as skeleton-based datasets. This paper proposes a comprehensive data augmentation framework that integrates geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Nada Aboudeshish , Dmitry Ignatov , Radu Timofte

Medical imaging is critical for diagnostics, but clinical adoption of advanced AI-driven imaging faces challenges due to patient variability, image artifacts, and limited model generalization. While deep learning has transformed image…

Image and Video Processing · Electrical Eng. & Systems 2025-06-02 Abdul-mojeed Olabisi Ilyas , Adeleke Maradesa , Jamal Banzi , Jianpan Huang , Henry K. F. Mak , Kannie W. Y. Chan