中文
相关论文

相关论文: Condition Weaving Meets Expert Modulation: Towards…

200 篇论文

Unified multimodal models hold the promise of generating extensive, interleaved narratives, weaving text and imagery into coherent long-form stories. However, current systems suffer from a critical reliability gap: as sequences grow,…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Haoyu Chen , Qing Liu , Yuqian Zhou , He Zhang , Zhaowen Wang , Mengwei Ren , Jingjing Ren , Xiang Wang , Zhe Lin , Lei Zhu

Dynamical systems in the life sciences are often composed of complex mixtures of overlapping behavioral regimes. Cellular subpopulations may shift from cycling to equilibrium dynamics or branch towards different developmental fates. The…

机器学习 · 计算机科学 2025-10-13 Nathan Quiblier , Roy Friedman , Matthew Ricci

Mixture-of-Experts (MoE) enhances model performance while maintaining computational efficiency, making it well-suited for large-scale applications. Conventional mixture-of-experts (MoE) architectures suffer from suboptimal coordination…

机器学习 · 计算机科学 2025-09-24 Yujiao Yang , Jing Lian , Linhui Li

Recent advances in generative artificial intelligence have had a significant impact on diverse domains spanning computer vision, natural language processing, and drug discovery. This work extends the reach of generative models into physical…

机器学习 · 计算机科学 2024-10-22 Christian Jacobsen , Yilin Zhuang , Karthik Duraisamy

Unsupervised image-to-image translation aims at learning the relationship between samples from two image domains without supervised pair information. The relationship between two domain images can be one-to-one, one-to-many or many-to-many.…

计算机视觉与模式识别 · 计算机科学 2018-05-21 Yongqi Zhang

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Rui Tian , Mingfei Gao , Haiming Gang , Jiasen Lu , Zhe Gan , Yinfei Yang , Zuxuan Wu , Afshin Dehghan

Large Language Models (LLMs) based on Mixture-of-Experts (MoE) are pivotal in industrial applications for their ability to scale performance efficiently. However, standard MoEs enforce uniform expert sizes,creating a rigidity that fails to…

计算与语言 · 计算机科学 2026-04-29 Zhicheng Ma , Xiang Liu , Zhaoxiang Liu , Ning Wang , Yi Shen , Kai Wang , Shuming Shi , Shiguo Lian

In this paper, we propose a unified layout planning and image generation model, PlanGen, which can pre-plan spatial layout conditions before generating images. Unlike previous diffusion-based models that treat layout planning and…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Runze He , Bo Cheng , Yuhang Ma , Qingxiang Jia , Shanyuan Liu , Ao Ma , Xiaoyu Wu , Liebucha Wu , Dawei Leng , Yuhui Yin

The solution of high-dimensional inference and prediction problems in computational biology is almost always a compromise between mathematical theory and practical constraints such as limited computational resources. As time progresses,…

定量方法 · 定量生物学 2009-01-13 Anagha Joshi , Riet De Smet , Kathleen Marchal , Yves Van de Peer , Tom Michoel

Real-world image degradation is often unknown, spatially non-uniform, and compositional, requiring all-in-one restoration models to adapt a single set of weights to diverse local corruption patterns without test-time degradation labels.…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Haisen He , Xiangyu Zou , SongLin Dong , Heng Li , Yihong Gong , Zhiheng Ma

Multimodal Variational Autoencoders have emerged as a popular tool to extract effective representations from rich multimodal data. However, such models rely on fusion strategies in latent space that destroy the joint statistical structure…

机器学习 · 计算机科学 2026-03-03 Federico Caretti , Guido Sanguinetti

Recently, diffusion models have been used successfully to fit distributions for cross-modal data translation and multimodal data generation. However, these methods rely on extensive scaling, overlooking the inefficiency and interference…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Zizhao Hu , Shaochong Jia , Mohammad Rostami

We introduce Skywork UniPic, a 1.5 billion-parameter autoregressive model that unifies image understanding, text-to-image generation, and image editing within a single architecture-eliminating the need for task-specific adapters or…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Peiyu Wang , Yi Peng , Yimeng Gan , Liang Hu , Tianyidan Xie , Xiaokun Wang , Yichen Wei , Chuanxin Tang , Bo Zhu , Changshi Li , Hongyang Wei , Eric Li , Xuchen Song , Yang Liu , Yahui Zhou

Although recent complex scene conditional generation models generate increasingly appealing scenes, it is very hard to assess which models perform better and why. This is often due to models being trained to fit different data splits, and…

计算机视觉与模式识别 · 计算机科学 2020-12-09 Arantxa Casanova , Michal Drozdzal , Adriana Romero-Soriano

This dissertation attempts to drive innovation in the field of generative modeling for computer vision, by exploring novel formulations of conditional generative models, and innovative applications in images, 3D animations, and video. Our…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Vikram Voleti

In recent years, Generative Adversarial Networks (GANs) have improved steadily towards generating increasingly impressive real-world images. It is useful to steer the image generation process for purposes such as content creation. This can…

计算机视觉与模式识别 · 计算机科学 2020-05-12 David Stap , Maurits Bleeker , Sarah Ibrahimi , Maartje ter Hoeve

Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models…

计算与语言 · 计算机科学 2025-06-02 Jiwan Chung , Janghan Yoon , Junhyeong Park , Sangeyl Lee , Joowon Yang , Sooyeon Park , Youngjae Yu

Image demoir\'eing poses one of the most formidable challenges in image restoration, primarily due to the unpredictable and anisotropic nature of moir\'e patterns. Limited by the quantity and diversity of training data, current methods tend…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Zemin Yang , Yujing Sun , Xidong Peng , Siu Ming Yiu , Yuexin Ma

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart,…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Xuanke Shi , Boxuan Li , Xiaoyang Han , Zhongang Cai , Lei Yang , Quan Wang , Dahua Lin

Multimodal image-to-image translation (I2IT) aims to learn a conditional distribution that explores multiple possible images in the target domain given an input image in the source domain. Conditional generative adversarial networks (cGANs)…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Zhiwen Zuo , Lei Zhao , Zhizhong Wang , Haibo Chen , Ailin Li , Qijiang Xu , Wei Xing , Dongming Lu