English
Related papers

Related papers: OmniColor: A Unified Framework for Multi-modal Lin…

200 papers

Modeling scenes using video generation models has garnered growing research interest in recent years. However, most existing approaches rely on perspective video models that synthesize only limited observations of a scene, leading to issues…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Yuheng Liu , Xin Lin , Xinke Li , Baihan Yang , Chen Wang , Kalyan Sunkavalli , Yannick Hold-Geoffroy , Hao Tan , Kai Zhang , Xiaohui Xie , Zifan Shi , Yiwei Hu

Representation learning, a task of learning latent vectors to represent entities, is a key task in improving search and recommender systems in web applications. Various representation learning methods have been developed, including…

Information Retrieval · Computer Science 2025-06-13 Anirudhan Badrinath , Alex Yang , Kousik Rajesh , Prabhat Agarwal , Jaewon Yang , Haoyu Chen , Jiajing Xu , Charles Rosenberg

Change detection (CD) in remote sensing is vital for applications such as urban monitoring and disaster assessment, yet traditional methods struggle with generalization across diverse scenarios. We present OmniCD, a foundational framework…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chenhao Sun

We present a fully automatic approach to video colorization with self-regularization and diversity. Our model contains a colorization network for video frame colorization and a refinement network for spatiotemporal color refinement. Without…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Chenyang Lei , Qifeng Chen

Object counting is pivotal for understanding the composition of scenes. Previously, this task was dominated by class-specific methods, which have gradually evolved into more adaptable class-agnostic strategies. However, these strategies…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Anindya Mondal , Sauradip Nag , Xiatian Zhu , Anjan Dutta

In this paper, we formulate the colorization problem into a multinomial classification problem and then apply a weighted function to classes. We propose a set of formulas to transform color values into color classes and vice versa. To…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Mrityunjoy Gain , Avi Deb Raha , Rameswar Debnath

We present MGM-Omni, a unified Omni LLM for omni-modal understanding and expressive, long-horizon speech generation. Unlike cascaded pipelines that isolate speech synthesis, MGM-Omni adopts a "brain-mouth" design with a dual-track,…

Sound · Computer Science 2025-09-30 Chengyao Wang , Zhisheng Zhong , Bohao Peng , Senqiao Yang , Yuqi Liu , Haokun Gui , Bin Xia , Jingyao Li , Bei Yu , Jiaya Jia

Unsupervised multiplex graph learning (UMGL) has been shown to achieve significant effectiveness for different downstream tasks by exploring both complementary information and consistent information among multiple graphs. However, previous…

Machine Learning · Computer Science 2023-08-04 Liang Peng , Xin Wang , Xiaofeng Zhu

Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms multi-task quality reasoning into continuous and interpretable…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yiting Lu , Fengbin Guan , Yixin Gao , Yan Zhong , Xinge Peng , Jiakang Yuan , Yihao Liu , Bo Zhang , Xin Li , Zhibo Chen , Weisi Lin

Colorizing grayscale images offers an engaging visual experience. Existing automatic colorization methods often fail to generate satisfactory results due to incorrect semantic colors and unsaturated colors. In this work, we propose an…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Han Wang , Xinning Chai , Yiwen Wang , Yuhong Zhang , Rong Xie , Li Song

Exemplar-based image colorization aims to colorize a target grayscale image based on a color reference image, and the key is to establish accurate pixel-level semantic correspondence between these two images. Previous methods search for…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Siqi Chen , Xueming Li , Xianlin Zhang , Mingdao Wang , Yu Zhang , Yue Zhang

Grayscale image colorization is a fascinating application of AI for information restoration. The inherently ill-posed nature of the problem makes it even more challenging since the outputs could be multi-modal. The learning-based methods…

Computer Vision and Pattern Recognition · Computer Science 2022-08-18 Himanshu Kumar , Abeer Banerjee , Sumeet Saurav , Sanjay Singh

Diffusion models have advanced image stylization significantly, yet two core challenges persist: (1) maintaining consistent stylization in complex scenes, particularly identity, composition, and fine details, and (2) preventing style…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Yiren Song , Cheng Liu , Mike Zheng Shou

The image matching field has been witnessing a continuous emergence of novel learnable feature matching techniques, with ever-improving performance on conventional benchmarks. However, our investigation shows that despite these gains, their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hanwen Jiang , Arjun Karpur , Bingyi Cao , Qixing Huang , Andre Araujo

Multimodal large language models (MLLMs) have shown strong capabilities but remain limited to fixed modality pairs and require costly fine-tuning with large aligned datasets. Building fully omni-capable models that can integrate text,…

Artificial Intelligence · Computer Science 2025-11-06 Huawei Lin , Yunzhi Shi , Tong Geng , Weijie Zhao , Wei Wang , Ravender Pal Singh

In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry.…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Zhenhuan Liu , Liang Li , Huajie Jiang , Xin Jin , Dandan Tu , Shuhui Wang , Zheng-Jun Zha

We present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Yiming Xie , Varun Jampani , Lei Zhong , Deqing Sun , Huaizu Jiang

Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intrinsic heterogeneity of diverse modalities presents substantial challenges to achieve…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Chengxuan Qian , Shuo Xing , Shawn Li , Yue Zhao , Zhengzhong Tu

In the practical application of restoring low-resolution gray-scale images, we generally need to run three separate processes of image colorization, super-resolution, and dows-sampling operation for the target device. However, this pipeline…

Computer Vision and Pattern Recognition · Computer Science 2022-01-13 Jiangning Zhang , Chao Xu , Jian Li , Yue Han , Yabiao Wang , Ying Tai , Yong Liu

Multimodal learning seeks to integrate information from heterogeneous sources, where signals may be shared across modalities, specific to individual modalities, or emerge only through their interaction. While self-supervised multimodal…

Machine Learning · Computer Science 2026-02-17 Carolin Cissee , Raneen Younis , Zahra Ahmadi
‹ Prev 1 4 5 6 7 8 10 Next ›