English
Related papers

Related papers: Frequency-Aware Flow Matching for High-Quality Ima…

200 papers

Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths. Both endpoints, however, concentrate in thin spherical shells, and a Euclidean chord leaves those shells even…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Tuna Han Salih Meral , Kaan Oktay , Hidir Yesiltepe , Adil Kaan Akan , Pinar Yanardag

Video Diffusion Models (VDMs) can generate high-quality videos, but often struggle with producing temporally coherent motion. Optical flow supervision is a promising approach to address this, with prior works commonly employing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Kuanting Wu , Kei Ota , Asako Kanezaki

Speech enhancement involves the distinction of a target speech signal from an intrusive background. Although generative approaches using Variational Autoencoders or Generative Adversarial Networks (GANs) have increasingly been used in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Martin Strauss , Bernd Edler

Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each training sample supervises only a single trajectory and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 George Stoica , Sayak Paul , Matthew Wallingford , Vivek Ramanujan , Abhay Nori , Winson Han , Ali Farhadi , Ranjay Krishna , Judy Hoffman

Flow Matching and Transformer architectures have demonstrated remarkable performance in image generation tasks, with recent work FlowAR [Ren et al., 2024] synergistically integrating both paradigms to advance synthesis fidelity. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Yingyu Liang , Zhizhou Sha , Zhenmei Shi , Zhao Song , Mingda Wan

We present FlowLM, a flow matching language model transformed from pre-trained diffusion language models via efficient fine-tuning. By re-aligning the curved sampling trajectories of diffusion models into straight-line flows, FlowLM enables…

Computation and Language · Computer Science 2026-05-21 Runzhe Zhang , Letian Chen , Wenpeng Zhang , Zhouhan Lin , Peilin Zhao

Reconstructing natural images from functional magnetic resonance imaging (fMRI) data remains a core challenge in natural decoding due to the mismatch between the richness of visual stimuli and the noisy, low resolution nature of fMRI…

Image and Video Processing · Electrical Eng. & Systems 2025-09-03 Junliang Ye , Lei Wang , Md Zakir Hossain

Image inpainting plays a vital role in restoring missing image regions and supporting high-level vision tasks, but traditional methods struggle with complex textures and large occlusions. Although Transformer-based approaches have…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Sijin He , Guangfeng Lin , Tao Li , Yajun Chen

Gravitational lensing is one of the most powerful probes of dark matter, yet creating high-fidelity lensed images at scale remains a bottleneck. Existing tools rely on ray-tracing or forward-modeling pipelines that, while precise, are…

Instrumentation and Methods for Astrophysics · Physics 2025-11-17 Hamees Sayed , Pranath Reddy , Michael W. Toomey , Sergei Gleyzer

Predicting low-energy molecular conformations given a molecular graph is an important but challenging task in computational drug discovery. Existing state-of-the-art approaches either resort to large scale transformer-based models that…

Quantitative Methods · Quantitative Biology 2024-10-31 Majdi Hassan , Nikhil Shenoy , Jungyoon Lee , Hannes Stark , Stephan Thaler , Dominique Beaini

Generative models for image generation are now commonly used for a wide variety of applications, ranging from guided image generation for entertainment to solving inverse problems. Nonetheless, training a generator is a non-trivial feat…

Machine Learning · Computer Science 2025-03-07 Eldad Haber , Shadab Ahamed , Md. Shahriar Rahim Siddiqui , Niloufar Zakariaei , Moshe Eliasof

Generative modeling-based visuomotor policies have been widely adopted in robotic manipulation, attributed to their ability to model multimodal action distributions. However, the high inference cost of multi-step sampling limits its…

Robotics · Computer Science 2026-02-19 Yifei Su , Ning Liu , Dong Chen , Zhen Zhao , Kun Wu , Meng Li , Zhiyuan Xu , Zhengping Che , Jian Tang

Normalizing Flows (NFs) learn invertible mappings between the data and a Gaussian distribution. Prior works usually suffer from two limitations. First, they add random noise to training samples or VAE latents as data augmentation,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Qinyu Zhao , Guangting Zheng , Tao Yang , Rui Zhu , Xingjian Leng , Stephen Gould , Liang Zheng

We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-24 Dong Yang , Yiyi Cai , Yuki Saito , Lixu Wang , Hiroshi Saruwatari

Generating synthetic fake faces, known as pseudo-fake faces, is an effective way to improve the generalization of DeepFake detection. Existing methods typically generate these faces by blending real or fake faces in spatial domain. While…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Hanzhe Li , Jiaran Zhou , Yuezun Li , Baoyuan Wu , Bin Li , Junyu Dong

Estimating three-dimensional conformations of a molecular graph allows insight into the molecule's biological and chemical functions. Fast generation of valid conformations is thus central to molecular modeling. Recent advances in…

Machine Learning · Computer Science 2025-02-18 Sohil Atul Shah , Vladlen Koltun

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds…

Sound · Computer Science 2025-01-10 Yi Yuan , Xubo Liu , Haohe Liu , Mark D. Plumbley , Wenwu Wang

Although diffusion-based real-world image restoration (Real-IR) has achieved remarkable progress, efficiently leveraging ultra-large-scale pre-trained text-to-image (T2I) models and fully exploiting their potential remain significant…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Purui Bai , Junxian Duan , Pin Wang , Jinhua Hao , Ming Sun , Chao Zhou , Huaibo Huang

Autonomous driving requires reasoning about interactions with surrounding traffic. A prevailing approach is large-scale imitation learning on expert driving datasets, aimed at generalizing across diverse real-world scenarios. For online…

Diffusion models approximate the denoising distribution as a Gaussian and predict its mean, whereas flow matching models reparameterize the Gaussian mean as flow velocity. However, they underperform in few-step sampling due to…

Machine Learning · Computer Science 2025-09-03 Hansheng Chen , Kai Zhang , Hao Tan , Zexiang Xu , Fujun Luan , Leonidas Guibas , Gordon Wetzstein , Sai Bi