English
Related papers

Related papers: ViscoNet: Bridging and Harmonizing Visual and Text…

200 papers

Long-tailed multi-label visual recognition poses a significant challenge, as images typically contain multiple labels with highly imbalanced class distributions, leading to biased models that favor head classes while underperforming on tail…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Wei Tang , Zuo-Zheng Wang , Kun Zhang , Tong Wei , Min-Ling Zhang

Given an input video of a person and a new garment, the objective of this paper is to synthesize a new video where the person is wearing the specified garment while maintaining spatiotemporal consistency. Although significant advances have…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Hung Nguyen , Quang Qui-Vinh Nguyen , Khoi Nguyen , Rang Nguyen

Image style transfer models based on convolutional neural networks usually suffer from high temporal inconsistency when applied to videos. Some video style transfer models have been proposed to improve temporal consistency, yet they fail to…

Computer Vision and Pattern Recognition · Computer Science 2018-11-02 Chang Gao , Derun Gu , Fangjun Zhang , Yizhou Yu

We provide a two-way integration for the widely adopted ControlNet by integrating external condition generation algorithms into a single dense prediction method and incorporating its individually trained image generation processes into a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Yilin Wang , Haiyang Xu , Xiang Zhang , Zeyuan Chen , Zhizhou Sha , Zirui Wang , Zhuowen Tu

This paper proposes a novel, robust, and lightweight supervised Convolutional Neural Network (CNN)-based technique for frame identification and synchronization, designed to enhance short-link communication performance in a screen-to-camera…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Vaigai Nayaki Yokar , Hoa Le-Minh , Xicong Li , Wai Lok Woo , Luis Nero Alves , Stanislav Zvanovec , Tran The Son , Zabih Ghassemlooy

Multimodal remote sensing semantic segmentation enhances scene interpretation by exploiting complementary physical cues from heterogeneous data. Although pretrained Vision Foundation Models (VFMs) provide strong general-purpose…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Haocheng Li , Juepeng Zheng , Shuangxi Miao , Ruibo Lu , Guosheng Cai , Haohuan Fu , Jianxi Huang

Large-scale vision-language models (LVLMs) pretrained on massive image-text pairs have achieved remarkable success in visual representations. However, existing paradigms to transfer LVLMs to downstream tasks encounter two primary…

Multimedia · Computer Science 2023-08-01 Chunjin Yang , Fanman Meng , Shuai Chen , Mingyu Liu , Runtong Zhang

Text-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling between the current text encoder and image decoder in T2I…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Can Qin , Ning Yu , Chen Xing , Shu Zhang , Zeyuan Chen , Stefano Ermon , Yun Fu , Caiming Xiong , Ran Xu

The conditional text-to-image diffusion models have garnered significant attention in recent years. However, the precision of these models is often compromised mainly for two reasons, ambiguous condition input and inadequate condition…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Sicheng Li , Keqiang Sun , Zhixin Lai , Xiaoshi Wu , Feng Qiu , Haoran Xie , Kazunori Miyata , Hongsheng Li

In breast ultrasound images, precise lesion segmentation is essential for early diagnosis; however, low contrast, speckle noise, and unclear boundaries make this difficult. Even though deep learning models have demonstrated potential,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Prateek Singh , Moumita Dholey , P. K. Vinod

We propose the width-resolution mutual learning method (MutualNet) to train a network that is executable at dynamic resource constraints to achieve adaptive accuracy-efficiency trade-offs at runtime. Our method trains a cohort of…

Computer Vision and Pattern Recognition · Computer Science 2020-03-25 Taojiannan Yang , Sijie Zhu , Chen Chen , Shen Yan , Mi Zhang , Andrew Willis

Paired RGB-thermal data has shown significant utility across a range of applications, including image fusion, object tracking, and anomaly detection; however, its broader adoption is constrained by the limited availability of aligned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Tseten Sherpa , Sikandar Ali , Shubham Parab , Haoyun Feng , Matthew Dennis , Keenan Gibbons , Verrah Otiende , Geoffrey H. Siwo

Video style transfer is getting more attention in AI community for its numerous applications such as augmented reality and animation productions. Compared with traditional image style transfer, performing this task on video presents new…

Computer Vision and Pattern Recognition · Computer Science 2021-01-21 Yingying Deng , Fan Tang , Weiming Dong , Haibin Huang , Chongyang Ma , Changsheng Xu

Multimodal object detection leveraging RGB and Infrared (IR) images is pivotal for robust perception in all-weather scenarios. While recent adapter-based approaches efficiently transfer RGB-pretrained foundation models to this task, they…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Xiantai Xiang , Guangyao Zhou , Zixiao Wen , Wenshuai Li , Ben Niu , Feng Wang , Lijia Huang , Qiantong Wang , Yuhan Liu , Zongxu Pan , Yuxin Hu

Recently, a growing number of work design unsupervised paradigms for point cloud processing to alleviate the limitation of expensive manual annotation and poor transferability of supervised methods. Among them, CrossPoint follows the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Hongyu Sun , Yongcai Wang , Xudong Cai , Xuewei Bai , Deying Li

Convolutional neural networks (CNNs) trained on object recognition achieve high task performance but continue to exhibit vulnerability under a range of visual perturbations and out-of-domain images, when compared with biological vision.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Lucas Piper , Arlindo L. Oliveira , Tiago Marques

With the growing demand for real-time video enhancement in live applications, existing methods often struggle to balance speed and effective exposure control, particularly under uneven lighting. We introduce RRNet (Rendering Relighting…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Wenlong Yang , Canran Jin , Weihang Yuan , Chao Wang , Lifeng Sun

We propose a Coefficient-to-Basis Network (C2BNet), a novel framework for solving inverse problems within the operator learning paradigm. C2BNet efficiently adapts to different discretizations through fine-tuning, using a pre-trained model…

Machine Learning · Computer Science 2025-03-12 Zecheng Zhang , Hao Liu , Wenjing Liao , Guang Lin

Translating information between text and image is a fundamental problem in artificial intelligence that connects natural language processing and computer vision. In the past few years, performance in image caption generation has seen…

Computer Vision and Pattern Recognition · Computer Science 2017-06-06 Hao Dong , Jingqing Zhang , Douglas McIlwraith , Yike Guo

Synthesizing visually impressive images that seamlessly align both text prompts and specific artistic styles remains a significant challenge in Text-to-Image (T2I) diffusion models. This paper introduces StyleBlend, a method designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Zichong Chen , Shijin Wang , Yang Zhou