English
Related papers

Related papers: Minuet: A Diffusion Autoencoder for Compact Semant…

200 papers

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Yuying Ge , Yizhuo Li , Yixiao Ge , Ying Shan

The cosmic microwave background power spectra are a primary window into the early universe. However, achieving interpretable, likelihood-compatible compression and fast inference under weak model assumptions remains challenging. We propose…

Cosmology and Nongalactic Astrophysics · Physics 2025-11-03 Tian-Yang Sun , Tian-Nuo Li , He Wang , Jing-Fei Zhang , Xin Zhang

The autonomous car must recognize the driving environment quickly for safe driving. As the Light Detection And Range (LiDAR) sensor is widely used in the autonomous car, fast semantic segmentation of LiDAR point cloud, which is the…

Computer Vision and Pattern Recognition · Computer Science 2022-02-22 Jaehyun Park , Chansoo Kim , Kichun Jo

Semi-supervised medical image segmentation aims to leverage limited annotated data and rich unlabeled data to perform accurate segmentation. However, existing semi-supervised methods are highly dependent on the quality of self-generated…

Image and Video Processing · Electrical Eng. & Systems 2024-07-16 Xinyu Liu , Wuyang Li , Yixuan Yuan

With the onset of large-scale astronomical surveys capturing millions of images, there is an increasing need to develop fast and accurate deconvolution algorithms that generalize well to different images. A powerful and accessible…

Instrumentation and Methods for Astrophysics · Physics 2022-11-18 Utsav Akhaury , Jean-Luc Starck , Pascale Jablonka , Frédéric Courbin , Kevin Michalewicz

Large Multimodal Models (LMMs) uniformly perceive video frames, creating computational inefficiency for videos with inherently varying temporal information density. This paper present \textbf{Quicksviewer}, an LMM with new perceiving…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Ji Qi , Yuan Yao , Yushi Bai , Bin Xu , Juanzi Li , Zhiyuan Liu , Tat-Seng Chua

Weak gravitational lensing is a powerful cosmological probe, with non--Gaussian features potentially containing the majority of the information. We examine constraints on the parameter triplet $(\Omega_m,w,\sigma_8)$ from non-Gaussian…

Cosmology and Nongalactic Astrophysics · Physics 2015-05-19 Andrea Petri , Jia Liu , Zoltan Haiman , Morgan May , Lam Hui , Jan M. Kratochvil

Latent diffusion models (LDMs) dominate high-quality image generation, yet integrating representation learning with generative modeling remains a challenge. We introduce a novel generative image modeling framework that seamlessly bridges…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Theodoros Kouzelis , Efstathios Karypidis , Ioannis Kakogeorgiou , Spyros Gidaris , Nikos Komodakis

Interpreting horizon-scale black hole images currently relies on computationally intensive General Relativistic Ray Tracing (GRRT) simulations, which pose a significant bottleneck for rapid parameter exploration and high-precision tests of…

General Relativity and Quantum Cosmology · Physics 2026-03-16 Ao Liu , Xudong Zhang , Lin Ding , Cuihong Wen , Wentao Liu , Jieci Wang

Analyzing and visualizing scientific ensemble datasets with high dimensionality and complexity poses significant challenges. Dimensionality reduction techniques and autoencoders are powerful tools for extracting features, but they often…

Machine Learning · Computer Science 2026-01-19 Hamid Gadirov , Lennard Manuel , Steffen Frey

Ultra-faint dwarf galaxies, which can be detected as resolved satellite systems of the Milky Way, are critical to understanding galaxy formation, evolution, and the nature of dark matter, as they are the oldest, smallest, most metal-poor,…

This study presents Latent Diffusion Autoencoder (LDAE), a novel encoder-decoder diffusion-based framework for efficient and meaningful unsupervised learning in medical imaging, focusing on Alzheimer disease (AD) using brain MR from the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Gabriele Lozupone , Alessandro Bria , Francesco Fontanella , Frederick J. A. Meijer , Claudio De Stefano , Henkjan Huisman

We present validated and forward-modelled galaxy luminosity functions and photometric predictions for the Vera C. Rubin Observatory Legacy Survey of Space and Time using the ASTRID cosmological hydrodynamical simulation. Galaxy magnitudes…

Deep observations of the Universe, usually as a part of sky surveys, are one of the symbols of the modern astronomy because they can allow big collaborations, exploiting multiple facilities and shared knowledge. The new generation of…

Instrumentation and Methods for Astrophysics · Physics 2020-12-18 Elisa Portaluri , Valentina Viotto , Roberto Ragazzoni , Carmelo Arcidiacono , Maria Bergomi , Marco Dima , Davide Greggio , Jacopo Farinato , Demetrio Magrin

Dust is a major component of the interstellar medium. Through scattering, absorption and thermal re-emission, it can profoundly alter astrophysical observations. Models for dust composition and distribution are necessary to better…

Despite the great progress made by deep CNNs in image semantic segmentation, they typically require a large number of densely-annotated images for training and are difficult to generalize to unseen object categories. Few-shot segmentation…

Computer Vision and Pattern Recognition · Computer Science 2020-02-10 Kaixin Wang , Jun Hao Liew , Yingtian Zou , Daquan Zhou , Jiashi Feng

We employ self-supervised representation learning to distill information from 76 million galaxy images from the Dark Energy Spectroscopic Instrument Legacy Imaging Surveys' Data Release 9. Targeting the identification of new strong…

Instrumentation and Methods for Astrophysics · Physics 2022-06-23 George Stein , Jacqueline Blaum , Peter Harrington , Tomislav Medan , Zarija Lukic

We present DC-AE 1.5, a new family of deep compression autoencoders for high-resolution diffusion models. Increasing the autoencoder's latent channel number is a highly effective approach for improving its reconstruction quality. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Junyu Chen , Dongyun Zou , Wenkun He , Junsong Chen , Enze Xie , Song Han , Han Cai

This paper describes a serendipitous galaxy cluster survey that we plan to conduct with the XMM X-ray satellite. We have modeled the expected properties of such a survey for three different cosmological models, using an extended…

Astrophysics · Physics 2007-05-23 A. Kathy Romer , Pedro T. P. Viana , Andrew R. Liddle , Robert G. Mann

Recent video generation models largely rely on video autoencoders that compress pixel-space videos into latent representations. However, existing video autoencoders suffer from three major limitations: (1) fixed-rate compression that wastes…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Yao Teng , Minxuan Lin , Xian Liu , Shuai Wang , Xiao Yang , Xihui Liu