English
Related papers

Related papers: Skywork UniPic 3.0: Unified Multi-Image Compositio…

200 papers

Pre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts. However, ambiguities inherent in natural language and out-of-distribution effects make it hard to synthesize image styles, that…

Computer Vision and Pattern Recognition · Computer Science 2023-06-02 Kihyuk Sohn , Nataniel Ruiz , Kimin Lee , Daniel Castro Chin , Irina Blok , Huiwen Chang , Jarred Barber , Lu Jiang , Glenn Entis , Yuanzhen Li , Yuan Hao , Irfan Essa , Michael Rubinstein , Dilip Krishnan

Multispectral and Hyperspectral Image Fusion (MHIF) aims to reconstruct high-resolution images by integrating low-resolution hyperspectral images (LRHSI) and high-resolution multispectral images (HRMSI). However, existing methods face…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Baisong Li

The growing complexity and scale of visual model pre-training have made developing and deploying multi-task computer-aided diagnosis (CAD) systems increasingly challenging and resource-intensive. Furthermore, the medical imaging community…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Yitao Zhu , Yuan Yin , Zhenrong Shen , Zihao Zhao , Haiyu Song , Sheng Wang , Dinggang Shen , Qian Wang

Robotic assembly in architectural construction faces a persistent bottleneck: existing planners are either highly specialized, requiring prohibitive retraining for every new geometric design, or operationally inefficient, treating…

Machine Learning · Computer Science 2026-05-20 Shih-Yu Lai , Chia-Ching Yen , Yang-Ting Shen , Peter Yichen Chen , Yu-Lun Liu , Bing-Yu Chen

Recent advancements in vision-language pre-training via contrastive learning have significantly improved performance across computer vision tasks. However, in the medical domain, obtaining multimodal data is often costly and challenging due…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Ameera Bawazir , Kebin Wu , Wenbin Li

Multi-person motion prediction is a complex and emerging field with significant real-world applications. Current state-of-the-art methods typically adopt dual-path networks to separately modeling spatial features and temporal features.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Kehua Qu , Rui Ding , Jin Tang

Social media popularity prediction plays a crucial role in content optimization, marketing strategies, and user engagement enhancement across digital platforms. However, predicting post popularity remains challenging due to the complex…

Multimedia · Computer Science 2025-07-02 Liliang Ye , Yunyao Zhang , Yafeng Wu , Yi-Ping Phoebe Chen , Junqing Yu , Wei Yang , Zikai Song

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

The generative AI technology offers an increasing variety of tools for generating entirely synthetic images that are increasingly indistinguishable from real ones. Unlike methods that alter portions of an image, the creation of completely…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Manos Schinas , Symeon Papadopoulos

Current multi-modal image fusion methods typically rely on task-specific models, leading to high training costs and limited scalability. While generative methods provide a unified modeling perspective, they often suffer from slow inference…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Huayi Zhu , Xiu Shu , Youqiang Xiong , Qiao Liu , Rui Chen , Di Yuan , Xiaojun Chang , Zhenyu He

In the early stages of semiconductor equipment development, obtaining large quantities of raw optical images poses a significant challenge. This data scarcity hinder the advancement of AI-powered solutions in semiconductor manufacturing. To…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 ChunLiang Wu , Xiaochun Li

Image fusion, a fundamental low-level vision task, aims to integrate multiple image sequences into a single output while preserving as much information as possible from the input. However, existing methods face several significant…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Zihan Cao , Yu Zhong , Ziqi Wang , Liang-Jian Deng

Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR…

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counterpart,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xuanke Shi , Boxuan Li , Xiaoyang Han , Zhongang Cai , Lei Yang , Quan Wang , Dahua Lin

Image matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Jiangwei Ren , Xingyu Jiang , Zizhuo Li , Dingkang Liang , Xin Zhou , Xiang Bai

This paper presents a unified multimodal pre-trained model called N\"UWA that can generate new or manipulate existing visual data (i.e., images and videos) for various visual synthesis tasks. To cover language, image, and video at the same…

Computer Vision and Pattern Recognition · Computer Science 2021-11-25 Chenfei Wu , Jian Liang , Lei Ji , Fan Yang , Yuejian Fang , Daxin Jiang , Nan Duan

Hand motion plays a central role in human interaction, yet modeling realistic 4D hand motion (i.e., 3D hand pose sequences over time) remains challenging. Research in this area is typically divided into two tasks: (1) Estimation approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Zhihao Sun , Tong Wu , Ruirui Tu , Daoguo Dong , Zuxuan Wu

While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, open-source alternatives significantly lag behind. Most academic models remain heavily fragmented, and the few existing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Kaihang Pan , Qi Tian , Jianwei Zhang , Weijie Kong , Jiangfeng Xiong , Yanxin Long , Shixue Zhang , Haiyi Qiu , Tan Wang , Zheqi Lv , Yue Wu , Liefeng Bo , Siliang Tang , Zhao Zhong

We study instruction-based image editing under professional workflows and identify three persistent challenges: (i) editors often over-edit, modifying content beyond the user's intent; (ii) existing models are largely single-turn, while…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Ruijie Ye , Jiayi Zhang , Zhuoxin Liu , Zihao Zhu , Siyuan Yang , Li Li , Tianfu Fu , Franck Dernoncourt , Yue Zhao , Jiacheng Zhu , Ryan Rossi , Wenhao Chai , Zhengzhong Tu

Despite impressive progress in high-fidelity image synthesis, generative models still struggle with logic-intensive instruction following, exposing a persistent reasoning--execution gap. Meanwhile, closed-source systems (e.g., Nano Banana)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Sashuai Zhou , Qiang Zhou , Jijin Hu , Hanqing Yang , Yue Cao , Junpeng Ma , Yinchao Ma , Jun Song , Tiezheng Ge , Cheng Yu , Bo Zheng , Zhou Zhao
‹ Prev 1 8 9 10 Next ›