English
Related papers

Related papers: Diverse and Accurate Image Description Using a Var…

200 papers

Image synthesis under multi-modal priors is a useful and challenging task that has received increasing attention in recent years. A major challenge in using generative models to accomplish this task is the lack of paired data containing all…

Computer Vision and Pattern Recognition · Computer Science 2022-06-13 Nithin Gopalakrishnan Nair , Wele Gedara Chaminda Bandara , Vishal M Patel

Modern visual world modeling systems increasingly rely on high-capacity architectures and large-scale data to produce plausible motion, yet they often fail to preserve underlying 3D geometry or physically consistent camera dynamics. A key…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Andrew Bond , Ilkin Umut Melanlioglu , Erkut Erdem , Aykut Erdem

In this note we present a generative model of natural images consisting of a deep hierarchy of layers of latent random variables, each of which follows a new type of distribution that we call rectified Gaussian. These rectified Gaussian…

Machine Learning · Statistics 2016-03-01 Tim Salimans

Transductive methods always outperform inductive methods in few-shot image classification scenarios. However, the existing few-shot methods contain a latent condition: the number of samples in each class is the same, which may be…

Computer Vision and Pattern Recognition · Computer Science 2022-05-31 Zaiyun Yang

Current variational dialog models have employed pre-trained language models (PLMs) to parameterize the likelihood and posterior distributions. However, the Gaussian assumption made on the prior distribution is incompatible with these…

Computation and Language · Computer Science 2023-10-25 Tianyu Yang , Thy Thy Tran , Iryna Gurevych

Variational autoencoders are powerful algorithms for identifying dominant latent structure in a single dataset. In many applications, however, we are interested in modeling latent structure and variation that are enriched in a target…

Machine Learning · Computer Science 2019-02-14 Abubakar Abid , James Zou

This paper proposes a convolutional neural network that can fuse high-level prior for semantic image segmentation. Motivated by humans' vision recognition system, our key design is a three-layer generative structure consisting of high-level…

Computer Vision and Pattern Recognition · Computer Science 2015-11-24 Haitian Zheng , Yebin Liu , Mengqi Ji , Feng Wu , Lu Fang

Variational auto-encoders are powerful probabilistic models in generative tasks but suffer from generating low-quality samples which are caused by the holes in the prior. We propose the Coupled Variational Auto-Encoder (C-VAE), which…

Machine Learning · Statistics 2023-06-06 Xiaoran Hao , Patrick Shafto

Significant progress has been made on visual captioning, largely relying on pre-trained features and later fixed object detectors that serve as rich inputs to auto-regressive models. A key limitation of such methods, however, is that the…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Chia-Wen Kuo , Zsolt Kira

We present a new method for improving the performances of variational autoencoder (VAE). In addition to enforcing the deep feature consistent principle thus ensuring the VAE output and its corresponding input images to have similar deep…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Xianxu Hou , Ke Sun , Linlin Shen , Guoping Qiu

Gaussian processes (GPs), implemented through multivariate Gaussian distributions for a finite collection of data, are the most popular approach in small-area spatial statistical modelling. In this context they are used to encode…

Machine Learning · Computer Science 2023-04-11 Elizaveta Semenova , Yidan Xu , Adam Howes , Theo Rashid , Samir Bhatt , Swapnil Mishra , Seth Flaxman

In this paper, we propose Language-Guided Contrastive Audio-Visual Masked Autoencoders (LG-CAV-MAE) to improve audio-visual representation learning. LG-CAV-MAE integrates a pretrained text encoder into contrastive audio-visual masked…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Yuchi Ishikawa , Shota Nakada , Hokuto Munakata , Kazuhiro Saito , Tatsuya Komatsu , Yoshimitsu Aoki

Large climate-model ensembles are computationally expensive; yet many downstream analyses would benefit from additional, statistically consistent realizations of spatiotemporal climate variables. We study a generative modeling approach for…

Machine Learning · Computer Science 2026-01-06 Jacquelyn Shelton , Przemyslaw Polewski , Alexander Robel , Matthew Hoffman , Stephen Price

Visual captioning aims to generate textual descriptions given images or videos. Traditionally, image captioning models are trained on human annotated datasets such as Flickr30k and MS-COCO, which are limited in size and diversity. This…

Computer Vision and Pattern Recognition · Computer Science 2021-03-01 Marimuthu Kalimuthu , Aditya Mogadala , Marius Mosbach , Dietrich Klakow

With the development of deep learning techniques, the combination of deep learning with image compression has drawn lots of attention. Recently, learned image compression methods had exceeded their classical counterparts in terms of…

Image and Video Processing · Electrical Eng. & Systems 2022-08-03 Ze Cui , Jing Wang , Shangyin Gao , Bo Bai , Tiansheng Guo , Yihui Feng

We propose a novel algorithm for quantizing continuous latent representations in trained models. Our approach applies to deep probabilistic models, such as variational autoencoders (VAEs), and enables both data and model compression. Unlike…

Image and Video Processing · Electrical Eng. & Systems 2020-09-09 Yibo Yang , Robert Bamler , Stephan Mandt

Layout generation is a novel task in computer vision, which combines the challenges in both object localization and aesthetic appraisal, widely used in advertisements, posters, and slides design. An accurate and pleasant layout should…

Computer Vision and Pattern Recognition · Computer Science 2022-09-05 Yunning Cao , Ye Ma , Min Zhou , Chuanbin Liu , Hongtao Xie , Tiezheng Ge , Yuning Jiang

This tutorial focuses on the fundamental architectures of Variational Autoencoders (VAE) and Generative Adversarial Networks (GAN), disregarding their numerous variations, to highlight their core principles. Both VAE and GAN utilize simple…

Machine Learning · Computer Science 2025-03-05 Yuan-Hao Wei

This paper aims to conduct a comparative analysis of contemporary Variational Autoencoder (VAE) architectures employed in anomaly detection, elucidating their performance and behavioral characteristics within this specific task. The…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Huy Hoang Nguyen , Cuong Nhat Nguyen , Xuan Tung Dao , Quoc Trung Duong , Dzung Pham Thi Kim , Minh-Tan Pham

Multi-label classification (MLC) is a prediction task where each sample can have more than one label. We propose a novel contrastive learning boosted multi-label prediction model based on a Gaussian mixture variational autoencoder…

Machine Learning · Computer Science 2022-06-13 Junwen Bai , Shufeng Kong , Carla P. Gomes
‹ Prev 1 4 5 6 7 8 10 Next ›