中文
相关论文

相关论文: Beyond Generation: Unlocking Universal Editing via…

200 篇论文

The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Yaofang Liu , Xiaodong Cun , Xuebo Liu , Xintao Wang , Yong Zhang , Haoxin Chen , Yang Liu , Tieyong Zeng , Raymond Chan , Ying Shan

Video generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Dacheng Li , Yunhao Fang , Yukang Chen , Shuo Yang , Shiyi Cao , Justin Wong , Michael Luo , Xiaolong Wang , Hongxu Yin , Joseph E. Gonzalez , Ion Stoica , Song Han , Yao Lu

Recent advances in unified multimodal models (UMMs) have enabled impressive progress in visual comprehension and generation. However, existing datasets and benchmarks focus primarily on single-turn interactions, failing to capture the…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Wei Chow , Jiachun Pan , Yongyuan Liang , Mingze Zhou , Xue Song , Liyu Jia , Saining Zhang , Siliang Tang , Juncheng Li , Fengda Zhang , Weijia Wu , Hanwang Zhang , Tat-Seng Chua

We consider the targeted image editing problem: blending a region in a source image with a driver image that specifies the desired change. Differently from prior works, we solve this problem by learning a conditional probability…

计算机视觉与模式识别 · 计算机科学 2022-05-04 Andrew Brown , Cheng-Yang Fu , Omkar Parkhi , Tamara L. Berg , Andrea Vedaldi

CLIP has become a cornerstone of multimodal representation learning, yet improving its performance typically requires a prohibitively costly process of training from scratch on billions of samples. We ask a different question: Can we…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Anant Mehta , Xiyuan Wei , Xingyu Chen , Tianbao Yang

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but they still face challenges in General Visual Editing, particularly in following complex instructions, preserving appearance…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Xiangyu Zhao , Peiyuan Zhang , Kexian Tang , Xiaorong Zhu , Hao Li , Wenhao Chai , Zicheng Zhang , Renqiu Xia , Guangtao Zhai , Junchi Yan , Hua Yang , Xue Yang , Haodong Duan

Unifying diverse image generation tasks within a single framework remains a fundamental challenge in visual generation. While large language models (LLMs) achieve unification through task-agnostic data and generation, existing visual…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Yijing Lin , Mengqi Huang , Shuhan Zhuang , Zhendong Mao

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we introduce VEGGIE, a…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Shoubin Yu , Difan Liu , Ziqiao Ma , Yicong Hong , Yang Zhou , Hao Tan , Joyce Chai , Mohit Bansal

3D editing refers to the ability to apply local or global modifications to 3D assets. Effective 3D editing requires maintaining semantic consistency by performing localized changes according to prompts, while also preserving local…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Yizhao Xu , Hongyuan Zhu , Caiyun Liu , Tianfu Wang , Keyu Chen , Sicheng Xu , Jiaolong Yang , Nicholas Jing Yuan , Qi Zhang

Pre-training general-purpose visual features with convolutional neural networks without relying on annotations is a challenging and important task. Most recent efforts in unsupervised feature learning have focused on either small or highly…

计算机视觉与模式识别 · 计算机科学 2019-08-14 Mathilde Caron , Piotr Bojanowski , Julien Mairal , Armand Joulin

Recent advances in deep generative models have demonstrated impressive results in photo-realistic facial image synthesis and editing. Facial expressions are inherently the result of muscle movement. However, existing neural network-based…

计算机视觉与模式识别 · 计算机科学 2019-11-07 ShahRukh Athar , Zhixin Shu , Dimitris Samaras

Underwater video enhancement (UVE) aims to improve the visibility and frame quality of underwater videos, which has significant implications for marine research and exploration. However, existing methods primarily focus on developing image…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Dazhao Du , Enhan Li , Lingyu Si , Fanjiang Xu , Jianwei Niu

The diffusion-based generative models have achieved remarkable success in text-based image generation. However, since it contains enormous randomness in generation progress, it is still challenging to apply such models for real-world visual…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Chenyang Qi , Xiaodong Cun , Yong Zhang , Chenyang Lei , Xintao Wang , Ying Shan , Qifeng Chen

Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets with precise spatial locations and temporal boundaries are…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Yiqing Shen , Chenjia Li , Mathias Unberath

Building upon large language models (LLMs), recent large multimodal models (LMMs) unify cross-model understanding and generation into a single framework. However, LMMs still struggle to achieve accurate vision-language alignment, prone to…

人工智能 · 计算机科学 2025-09-09 Jixiang Hong , Yiran Zhang , Guanzhong Wang , Yi Liu , Ji-Rong Wen , Rui Yan

Recent video foundation models such as SAM2 excel at prompted video segmentation by treating masks as a general-purpose primitive. However, many real-world settings require unprompted segmentation that aims to detect and track all objects…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Miran Heo , Sukjun Hwang , Min-Hung Chen , Yu-Chiang Frank Wang , Albert Gu , Seon Joo Kim , Ryo Hachiuma

We present IMAS, a method that segments the primary objects in videos without manual annotation in training or inference. Previous methods in unsupervised video object segmentation (UVOS) have demonstrated the effectiveness of motion as…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Long Lian , Zhirong Wu , Stella X. Yu

Unified Multimodal Models (UMMs) have emerged as a promising paradigm that integrates multimodal understanding and generation within a unified modeling framework. However, current generative training paradigms suffer from inherent…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Jiyeong Kim , Yerim So , Hyesong Choi , Uiwon Hwang , Dongbo Min

Unsupervised pre-training can equip reinforcement learning agents with prior knowledge and accelerate learning in downstream tasks. A promising direction, grounded in human development, investigates agents that learn by setting and pursuing…

机器学习 · 计算机科学 2026-01-28 Octavio Pappalardo

We introduce Generative Universal Verifier, a novel concept and plugin designed for next-generation multimodal reasoning in vision-language models and unified multimodal models, providing the fundamental capability of reflection and…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Xinchen Zhang , Xiaoying Zhang , Youbin Wu , Yanbin Cao , Renrui Zhang , Ruihang Chu , Ling Yang , Yujiu Yang