English
Related papers

Related papers: Multi-scale Latent Point Consistency Models for 3D…

200 papers

This technical report introduces PIXART-{\delta}, a text-to-image synthesis framework that integrates the Latent Consistency Model (LCM) and ControlNet into the advanced PIXART-{\alpha} model. PIXART-{\alpha} is recognized for its ability…

Computer Vision and Pattern Recognition · Computer Science 2024-01-11 Junsong Chen , Yue Wu , Simian Luo , Enze Xie , Sayak Paul , Ping Luo , Hang Zhao , Zhenguo Li

Estimating reliable geometric model parameters from the data with severe outliers is a fundamental and important task in computer vision. This paper attempts to sample high-quality subsets and select model instances to estimate parameters…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Guobao Xiao , Jun Yu , Jiayi Ma , Deng-Ping Fan , Ling Shao

Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent…

Computation and Language · Computer Science 2026-05-11 Viacheslav Meshchaninov , Alexander Shabalin , Egor Chimbulatov , Nikita Gushchin , Ilya Koziev , Alexander Korotin , Dmitry Vetrov

High-resolution 3D object generation remains a challenging task primarily due to the limited availability of comprehensive annotated training data. Recent advancements have aimed to overcome this constraint by harnessing image generative…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Zijie Pan , Jiachen Lu , Xiatian Zhu , Li Zhang

Computer System Architecture serves as a crucial bridge between software applications and the underlying hardware, encompassing components like compilers, CPUs, coprocessors, and RTL designs. Its development, from early mainframes to modern…

Most music generation models directly generate a single music mixture. To allow for more flexible and controllable generation, the Multi-Source Diffusion Model (MSDM) has been proposed to model music as a mixture of multiple instrumental…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-18 Zhongweiyang Xu , Debottam Dutta , Yu-Lin Wei , Romit Roy Choudhury

Light-field microscopy (LFM) enables single-shot capture of multi-angular information from biological samples, supporting real-time volumetric imaging. However, traditional physics-based algorithms often suffer from limited spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Qihong Zhao , Shaokang Yan , Zhimin Qiao , Jinjia Wang , Bo Xiong

In medical image segmentation tasks, diffusion models have shown significant potential. However, mainstream diffusion models suffer from drawbacks such as multiple sampling times and slow prediction results. Recently, consistency models, as…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Kejia Zhang , Lan Zhang , Haiwei Pan , Baolong Yu

Multi-modal images play a crucial role in comprehensive evaluations in medical image analysis providing complementary information for identifying clinically important biomarkers. However, in clinical practice, acquiring multiple modalities…

Image and Video Processing · Electrical Eng. & Systems 2024-10-02 Jonghun Kim , Hyunjin Park

Latent Diffusion Models (LDMs) produce high-quality, photo-realistic images, however, the latency incurred by multiple costly inference iterations can restrict their applicability. We introduce LatentCRF, a continuous Conditional Random…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Kanchana Ranasinghe , Sadeep Jayasumana , Andreas Veit , Ayan Chakrabarti , Daniel Glasner , Michael S Ryoo , Srikumar Ramalingam , Sanjiv Kumar

Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. However, these models are far from real-time, often requiring…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Hanzhong Guo , Hongwei Yi , Daquan Zhou , Alexander William Bergman , Michael Lingelbach , Yizhou Yu

Textile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated TPG, existing image-to-image models appear to be natural…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Chenggong Hu , Yi Wang , Mengqi Xue , Haofei Zhang , Jie Song , Li Sun

Many particle physics datasets like those generated at colliders are described by continuous coordinates (in contrast to grid points like in an image), respect a number of symmetries (like permutation invariance), and have a stochastic…

High Energy Physics - Phenomenology · Physics 2023-11-03 Vinicius Mikuni , Benjamin Nachman , Mariel Pettee

Diffusion Probabilistic Models (DPMs) suffer from inefficient inference due to their slow sampling and high memory consumption, which limits their applicability to various medical imaging applications. In this work, we propose a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Fahim Ahmed Zaman , Mathews Jacob , Amanda Chang , Kan Liu , Milan Sonka , Xiaodong Wu

Diffusion models are the de facto approach for generating high-quality images and videos, but learning high-dimensional models remains a formidable task due to computational and optimization challenges. Existing methods often resort to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Jiatao Gu , Shuangfei Zhai , Yizhe Zhang , Josh Susskind , Navdeep Jaitly

We present PointInfinity, an efficient family of point cloud diffusion models. Our core idea is to use a transformer-based architecture with a fixed-size, resolution-invariant latent representation. This enables efficient training with…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Zixuan Huang , Justin Johnson , Shoubhik Debnath , James M. Rehg , Chao-Yuan Wu

In the medical domain, acquiring large datasets is challenging due to both accessibility issues and stringent privacy regulations. Consequently, data availability and privacy protection are major obstacles to applying machine learning in…

Image and Video Processing · Electrical Eng. & Systems 2025-07-02 Wenwu Tang , Khaled Seyam , Bin Yang

In this paper, we propose a novel garment-centric outpainting (GCO) framework based on the latent diffusion model (LDM) for fine-grained controllable apparel showcase image generation. The proposed framework aims at customizing a fashion…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Rong Zhang , Jingnan Wang , Zhiwen Zuo , Jianfeng Dong , Wei Li , Chi Wang , Weiwei Xu , Xun Wang

In this study we develop dimension-reduction techniques to accelerate diffusion model inference in the context of synthetic data generation. The idea is to integrate compressed sensing into diffusion models (hence, CSDM): First, compress…

Machine Learning · Statistics 2025-09-30 Zhengyi Guo , Jiatu Li , Wenpin Tang , David D. Yao

Non-contrast CT (NCCT) imaging may reduce image contrast and anatomical visibility, potentially increasing diagnostic uncertainty. In contrast, contrast-enhanced CT (CECT) facilitates the observation of regions of interest (ROI). Leading…

Image and Video Processing · Electrical Eng. & Systems 2024-11-18 Tingyi Lin , Pengju Lyu , Jie Zhang , Yuqing Wang , Cheng Wang , Jianjun Zhu