English
Related papers

Related papers: Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image …

200 papers

Recent image generation models produce impressive composites, but often fail to preserve the identity of user-provided content when editing specific elements: the surrounding scene may shift, and even the edited object's appearance can…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Jinrui Yang , Qing Liu , Yijun Li , Mengwei Ren , Letian Zhang , Zhe Lin , Cihang Xie , Yuyin Zhou

Diffusion models have made significant progress in both text-to-image (T2I) generation and text-guided image editing. However, these models are typically built with billions of parameters, leading to high latency and increased deployment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Kailai Feng , Yuxiang Wei , Bo Chen , Yang Pan , Hu Ye , Songwei Liu , Chenqian Yan , Yuan Gao

Applying LLMs to complex industrial processes remains challenging due to the semantic gap between natural language design intents and the rigorous physical logic of engineering. In the field of petroleum refining engineering, a critical…

Computational Engineering, Finance, and Science · Computer Science 2026-05-20 Dongxiao Liu , Yuwen Ding , Xinghai Wei , Jiacheng Ji , Lei Li , Linghui Li , Xiaoyong Li

The rapid progress of large multimodal models has inspired efforts toward unified frameworks that couple understanding and generation. While such paradigms have shown remarkable success in 2D, extending them to 3D remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Yongwei Chen , Tianyi Wei , Yushi Lan , Zhaoyang Lyu , Shangchen Zhou , Xudong Xu , Xingang Pan

While latent diffusion models (LDMs), such as Stable Diffusion, are designed for high-resolution (HR) image generation, they often struggle with significant structural distortions when generating images at resolutions higher than their…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Boyuan Cao , Jiaxin Ye , Yujie Wei , Hongming Shan

The increased demand for tools that automate the 3D content creation process led to tremendous progress in deep generative models that can generate diverse 3D objects of high fidelity. In this paper, we present PASTA, an autoregressive…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Songlin Li , Despoina Paschalidou , Leonidas Guibas

We propose a method to fuse frozen text-only large language models (LLMs) with pre-trained image encoder and decoder models, by mapping between their embedding spaces. Our model demonstrates a wide suite of multimodal capabilities: image…

Computation and Language · Computer Science 2023-10-16 Jing Yu Koh , Daniel Fried , Ruslan Salakhutdinov

Text-guided image manipulation has experienced notable advancement in recent years. In order to mitigate linguistic ambiguity, few-shot learning with visual examples has been applied for instructions that are underrepresented in the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Bolin Lai , Felix Juefei-Xu , Miao Liu , Xiaoliang Dai , Nikhil Mehta , Chenguang Zhu , Zeyi Huang , James M. Rehg , Sangmin Lee , Ning Zhang , Tong Xiao

Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Yuxuan Xue , Ruofan Liang , Egor Zakharov , Timur Bagautdinov , Chen Cao , Giljoo Nam , Shunsuke Saito , Gerard Pons-Moll , Javier Romero

Surface cutting is a fundamental task in computer graphics, with applications in UV parameterization, texture mapping, and mesh decomposition. However, existing methods often produce technically valid but overly fragmented atlases that lack…

Real-world processes often generate data that are a mix of categorical and numeric values that are recorded at irregular and informative intervals. Discrete token-based approaches are limited in numeric representation capacity while methods…

Machine Learning · Computer Science 2025-06-02 Andrew J. Loza , Jun Yup Kim , Shangzheng Song , Yihang Liu , Joseph J. Y. Sung , R Andrew Taylor , Dennis L. Shung

The Swapping Autoencoder achieved state-of-the-art performance in deep image manipulation and image-to-image translation. We improve this work by introducing a simple yet effective auxiliary module based on gradient reversal layers. The…

Computer Vision and Pattern Recognition · Computer Science 2022-08-25 Shima Shahfar , Charalambos Poullis

Document reconstruction constitutes a significant facet of document analysis and recognition, a field that has been progressively accruing interest within the scholarly community. A multitude of these researchers employ an array of document…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Xin Li , Mingming Gong , Yunfei Wu , Jianxin Dai , Antai Guo , Xinghua Jiang , Haoyu Cao , Yinsong Liu , Deqiang Jiang , Xing Sun

Large-scale diffusion models have achieved remarkable success in generating high-quality images from textual descriptions, gaining popularity across various applications. However, the generation of layered content, such as transparent…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Yusuf Dalva , Yijun Li , Qing Liu , Nanxuan Zhao , Jianming Zhang , Zhe Lin , Pinar Yanardag

Depth acquisition, based on active illumination, is essential for autonomous and robotic navigation. LiDARs (Light Detection And Ranging) with mechanical, fixed, sampling templates are commonly used in today's autonomous vehicles. An…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Adam Wolff , Shachar Praisler , Ilya Tcenov , Guy Gilboa

Recent text-to-video (T2V) generation methods have seen significant advancements. However, the majority of these works focus on producing short video clips of a single event (i.e., single-scene videos). Meanwhile, recent large language…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Han Lin , Abhay Zala , Jaemin Cho , Mohit Bansal

Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understanding and generation remains challenging due to optimization…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Jie Qin , Jiancheng Huang , Limeng Qiao , Lin Ma

Large Language Models (LLMs) inherently use autoregressive decoding, which lacks parallelism in inference and results in significantly slow inference speed. While methods such as Medusa constructs parallelized heads, they lack adequate…

Artificial Intelligence · Computer Science 2024-10-21 Zeping Li , Xinlong Yang , Ziheng Gao , Ji Liu , Guanchen Li , Zhuang Liu , Dong Li , Jinzhang Peng , Lu Tian , Emad Barsoum

Transparency-aware generation requires modeling not only RGB appearance but also alpha-based opacity and cross-layer composition, which are essential for tasks such as image matting, object removal, layer decomposition, and multi-layer…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Hao Yu , Jinglin Wang , Jiabo Zhan , Rui Chen , Zile Wang , Huaisong Zhang , Hongyu Li , Xinrui Chen , Yongxian Wei , Chun Yuan

We present BLIP3o-NEXT, a fully open-source foundation model in the BLIP3 series that advances the next frontier of native image generation. BLIP3o-NEXT unifies text-to-image generation and image editing within a single architecture,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Jiuhai Chen , Le Xue , Zhiyang Xu , Xichen Pan , Shusheng Yang , Can Qin , An Yan , Honglu Zhou , Zeyuan Chen , Lifu Huang , Tianyi Zhou , Junnan Li , Silvio Savarese , Caiming Xiong , Ran Xu
‹ Prev 1 8 9 10 Next ›