English
Related papers

Related papers: Lavida-O: Elastic Large Masked Diffusion Models fo…

200 papers

Recent advances in generative modeling have positioned diffusion models as state-of-the-art tools for sampling from complex data distributions. While these models have shown remarkable success across single-modality domains such as images…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Nimrod Berman , Omkar Joglekar , Eitan Kosman , Dotan Di Castro , Omri Azencot

Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each problem setting, which fixes the input-output mapping and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Houyuan Chen , Hong Li , Xianghao Kong , Tianrui Zhu , Shaocong Xu , Weiqing Xiao , Yuwei Guo , Chongjie Ye , Lvmin Zhang , Hao Zhao , Anyi Rao

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MMGen, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Jiepeng Wang , Zhaoqing Wang , Hao Pan , Yuan Liu , Dongdong Yu , Changhu Wang , Wenping Wang

We introduce \textbf{Evo}, a duality latent trajectory model that bridges autoregressive (AR) and diffusion-based language generation within a continuous evolutionary generative framework. Rather than treating AR decoding and diffusion…

Machine Learning · Computer Science 2026-03-10 Junde Wu , Minhao Hu , Jiayuan Zhu , Yuyuan Liu , Tianyi Zhang , Kang Li , Jingkun Chen , Jiazhen Pan , Min Xu , Yueming Jin

Denoising diffusion probabilistic models have recently demonstrated state-of-the-art generative performance and have been used as strong pixel-level representation learners. This paper decomposes the interrelation between the generative…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Zixuan Pan , Jianxu Chen , Yiyu Shi

In emergencies, the ability to quickly and accurately gather environmental data and command information, and to make timely decisions, is particularly critical. Traditional semantic communication frameworks, primarily based on a single…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Weiqi Fu , Lianming Xu , Xin Wu , Haoyang Wei , Li Wang

This paper presents ThinkDiff, a novel alignment paradigm that empowers text-to-image diffusion models with multimodal in-context understanding and reasoning capabilities by integrating the strengths of vision-language models (VLMs).…

Machine Learning · Computer Science 2025-02-18 Zhenxing Mi , Kuan-Chieh Wang , Guocheng Qian , Hanrong Ye , Runtao Liu , Sergey Tulyakov , Kfir Aberman , Dan Xu

Masked diffusion models (MDM) are powerful generative models for discrete data that generate samples by progressively unmasking tokens in a sequence. Each token can take one of two states: masked or unmasked. We observe that token sequences…

Machine Learning · Computer Science 2025-10-23 Chen-Hao Chao , Wei-Fang Sun , Hanwen Liang , Chun-Yi Lee , Rahul G. Krishnan

While multimodal large language models (MLLMs) provide advanced reasoning for autonomous driving, translating their discrete semantic knowledge into continuous trajectories remains a fundamental challenge. Existing methods often rely on…

Robotics · Computer Science 2026-03-03 Fabian Schmidt , Karol Fedurko , Markus Enzweiler , Abhinav Valada

Autoregressive models (ARMs) have long dominated the landscape of biomedical vision-language models (VLMs). Recently, masked diffusion models such as LLaDA have emerged as promising alternatives, yet their application in the biomedical…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Xuanzhao Dong , Wenhui Zhu , Xiwen Chen , Zhipeng Wang , Peijie Qiu , Shao Tang , Xin Li , Yalin Wang

Visual counterfactual explanations aim to reveal the minimal semantic modifications that can alter a model's prediction, providing causal and interpretable insights into deep neural networks. However, existing diffusion-based counterfactual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Changlu Guo , Anders Nymark Christensen , Anders Bjorholm Dahl , Morten Rieger Hannemose

Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Tao Lin , Yuxin Du , Jiting Liu , Nuobei Zhu , Yunhe Li , Yuqian Fu , Yinxinyu Chen , Hongyi Cai , Zewei Ye , Bing Cheng , Kai Ye , Yiran Mao , Yilei Zhong , MingKang Dong , Junchi Yan , Gen Li , Bo Zhao

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models, these generations are not…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Hanzhe Hu , Zhizhuo Zhou , Varun Jampani , Shubham Tulsiani

This paper presents improved native unified multimodal models, \emph{i.e.,} Show-o2, that leverage autoregressive modeling and flow matching. Built upon a 3D causal variational autoencoder space, unified visual representations are…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Jinheng Xie , Zhenheng Yang , Mike Zheng Shou

The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Jinlong Li , Cristiano Saltori , Fabio Poiesi , Nicu Sebe

One-step generators distilled from Masked Diffusion Models (MDMs) compress multiple sampling steps into a single forward pass, enabling efficient text and image synthesis. However, they suffer two key limitations: they inherit modeling bias…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Yuanzhi Zhu , Xi Wang , Stéphane Lathuilière , Vicky Kalogeiton

Predictive manipulation has recently gained considerable attention in the Embodied AI community due to its potential to improve robot policy performance by leveraging predicted states. However, generating accurate future visual states of…

Robotics · Computer Science 2025-09-15 Yuhang Huang , Jiazhao Zhang , Shilong Zou , Xinwang Liu , Ruizhen Hu , Kai Xu

Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation methods may not generalize well to synthesized images or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yuanhao Zhai , Kevin Lin , Linjie Li , Chung-Ching Lin , Jianfeng Wang , Zhengyuan Yang , David Doermann , Junsong Yuan , Zicheng Liu , Lijuan Wang

3D molecule generation is crucial for drug discovery and material science, requiring models to process complex multi-modalities, including atom types, chemical bonds, and 3D coordinates. A key challenge is integrating these modalities of…

Machine Learning · Computer Science 2025-10-14 Yanchen Luo , Zhiyuan Liu , Yi Zhao , Sihang Li , Hengxing Cai , Kenji Kawaguchi , Tat-Seng Chua , Yang Zhang , Xiang Wang

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip the Diffusion…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Yangming Shi , Shixiang Zhu , Tao Shen , Zhimiao Yu , Dengsheng Chen , Taicai Chen , Yunfei Yang , Juan Zhou , Chen Cheng , Liang Ma , Xibin Wu , Benxuan Yan , Ge Li , Tuoyu Zhang , Dan Li , Chang Liu , Zhenbang Sun
‹ Prev 1 3 4 5 6 7 10 Next ›