English
Related papers

Related papers: OmniAlpha: Aligning Transparency-Aware Generation …

200 papers

Image matting refers to predicting the alpha values of unknown foreground areas from natural images. Prior methods have focused on propagating alpha values from known to unknown regions. However, not all natural images have a specifically…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Huanqia Cai , Fanglei Xue , Lele Xu , Lili Guo

The performance of robotic imitation learning is fundamentally limited by data quality and training strategies. Prevalent sampling strategies on RLBench suffer from severe keyframe redundancy and imbalanced temporal distribution, leading to…

Robotics · Computer Science 2026-03-03 Fanqi Pu , Lei Jiang , Wenming Yang

This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based visual encoder for both image and video inputs, and thus can…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Junke Wang , Dongdong Chen , Zuxuan Wu , Chong Luo , Luowei Zhou , Yucheng Zhao , Yujia Xie , Ce Liu , Yu-Gang Jiang , Lu Yuan

Unified image fusion aims to integrate complementary information from multi-source images, enhancing image quality through a unified framework applicable to diverse fusion tasks. While treating all fusion tasks as a unified problem…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 Xingyu Hu , Junjun Jiang , Chenyang Wang , Kui Jiang , Xianming Liu , Jiayi Ma

The pursuit of general-purpose artificial intelligence depends on large language models (LLMs) that can handle both structured reasoning and open-ended generation. We present Omni-Thinker, a unified reinforcement learning (RL) framework…

Machine Learning · Computer Science 2025-09-30 Derek Li , Jiaming Zhou , Leo Maxime Brunswic , Abbas Ghaddar , Qianyi Sun , Liheng Ma , Yu Luo , Dong Li , Mark Coates , Jianye Hao , Yingxue Zhang

We present UniGen-1.5, a unified multimodal large language model (MLLM) for advanced image understanding, generation and editing. Building upon UniGen, we comprehensively enhance the model architecture and training pipeline to strengthen…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Rui Tian , Mingfei Gao , Haiming Gang , Jiasen Lu , Zhe Gan , Yinfei Yang , Zuxuan Wu , Afshin Dehghan

Semantic analysis on visible (RGB) and infrared (IR) images has gained significant attention due to their enhanced accuracy and robustness under challenging conditions including low-illumination and adverse weather. However, due to the lack…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Maoxun Yuan , Bo Cui , Tianyi Zhao , Jiayi Wang , Shan Fu , Xue Yang , Xingxing Wei

We present Lunima-OmniLV (abbreviated as OmniLV), a universal multimodal multi-task framework for low-level vision that addresses over 100 sub-tasks across four major categories: image restoration, image enhancement, weak-semantic dense…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Yuandong Pu , Le Zhuo , Kaiwen Zhu , Liangbin Xie , Wenlong Zhang , Xiangyu Chen , Peng Gao , Yu Qiao , Chao Dong , Yihao Liu

We present an explainable, bias-aware generative framework that unifies cross-modal attention fusion, Grad-CAM++ attribution, and a Reveal-to-Revise feedback loop within a single training paradigm. The architecture couples a conditional…

Machine Learning · Computer Science 2026-04-08 Noor Islam S. Mohammad , Md Muntaqim Meherab

Visual anomaly detection aims to learn normality from normal images, but existing approaches are fragmented across various tasks: defect detection, semantic anomaly detection, multi-class anomaly detection, and anomaly clustering. This…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Yujin Lee , Harin Lim , Seoyoon Jang , Hyunsoo Yoon

The image matching field has been witnessing a continuous emergence of novel learnable feature matching techniques, with ever-improving performance on conventional benchmarks. However, our investigation shows that despite these gains, their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hanwen Jiang , Arjun Karpur , Bingyi Cao , Qixing Huang , Andre Araujo

Foundation models are multi-dataset and multi-task machine learning methods that once pre-trained can be fine-tuned for a large variety of downstream applications. The successful development of such general-purpose models for physics data…

High Energy Physics - Phenomenology · Physics 2024-09-10 Joschka Birk , Anna Hallin , Gregor Kasieczka

Multi-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative models. In this paper, we introduce the Anonymous Region…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Yifan Pu , Yiming Zhao , Zhicong Tang , Ruihong Yin , Haoxing Ye , Yuhui Yuan , Dong Chen , Jianmin Bao , Sirui Zhang , Yanbin Wang , Lin Liang , Lijuan Wang , Ji Li , Xiu Li , Zhouhui Lian , Gao Huang , Baining Guo

Diffusion models have recently motivated great success in many generation tasks like object removal. Nevertheless, existing image decomposition methods struggle to disentangle semi-transparent or transparent layer occlusions due to mask…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Zitong Wang , Hang Zhao , Qianyu Zhou , Xuequan Lu , Xiangtai Li , Yiren Song

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) offers a promising…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Guohui Zhang , XiaoXiao Ma , Jie Huang , Hang Xu , Hu Yu , Siming Fu , Yuming Li , Zeyue Xue , Lin Song , Haoyang Huang , Nan Duan , Feng Zhao

Transparent object perception remains a major challenge in computer vision research, as transparency confounds both depth estimation and semantic segmentation. Recent work has explored multi-task learning frameworks to improve robustness,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Gbenga Omotara , Ramy Farag , Seyed Mohamad Ali Tousi , G. N. DeSouza

Modern multimodal large language models (MLLMs) generate fluent responses from interleaved text, image, audio, and video inputs. However, identifying which input sources support each generated statement remains an open challenge. Existing…

Computation and Language · Computer Science 2026-04-16 Qianqi Yan , Yichen Guo , Ching-Chen Kuo , Shan Jiang , Hang Yin , Yang Zhao , Xin Eric Wang

Lineart colorization is a critical stage in professional content creation, yet achieving precise and flexible results under diverse user constraints remains a significant challenge. To address this, we propose OmniColor, a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Xulu Zhang , Haoqian Du , Xiaoyong Wei , Qing Li

We introduce OneDiffusion, a versatile, large-scale diffusion model that seamlessly supports bidirectional image synthesis and understanding across diverse tasks. It enables conditional generation from inputs such as text, depth, pose,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Duong H. Le , Tuan Pham , Sangho Lee , Christopher Clark , Aniruddha Kembhavi , Stephan Mandt , Ranjay Krishna , Jiasen Lu

Most previous image matting methods require a roughly-specificed trimap as input, and estimate fractional alpha values for all pixels that are in the unknown region of the trimap. In this paper, we argue that directly estimating the alpha…

Computer Vision and Pattern Recognition · Computer Science 2019-09-12 Shaofan Cai , Xiaoshuai Zhang , Haoqiang Fan , Haibin Huang , Jiangyu Liu , Jiaming Liu , Jiaying Liu , Jue Wang , Jian Sun