English
Related papers

Related papers: Disentangled Latent Energy-Based Style Translation…

200 papers

Tensor-based multi-view clustering has recently received significant attention due to its exceptional ability to explore cross-view high-order correlations. However, most existing methods still encounter some limitations. (1) Most of them…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Deng Xu , Chao Zhang , Zechao Li , Chunlin Chen , Huaxiong Li

Multilingual pre-trained models are able to zero-shot transfer knowledge from rich-resource to low-resource languages in machine reading comprehension (MRC). However, inherent linguistic discrepancies in different languages could make…

Computation and Language · Computer Science 2023-01-18 Linjuan Wu , Shaojuan Wu , Xiaowang Zhang , Deyi Xiong , Shizhan Chen , Zhiqiang Zhuang , Zhiyong Feng

This research introduces a transformative framework for integrating Vision-Enhanced Large Language Models (LLMs) with advanced transformer-based architectures to tackle challenges in high-resolution image synthesis and multimodal data…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Karthikeya KV

Recent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additionally, image and video tasks are typically handled in…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Huiang He , Minghui Hu , Chuanxia Zheng , Chaoyue Wang , Tat-Jen Cham

Accurate fMRI analysis requires sensitivity to temporal structure across multiple scales, as BOLD signals encode cognitive processes that emerge from fast transient dynamics to slower, large-scale fluctuations. Existing deep learning (DL)…

Signal Processing · Electrical Eng. & Systems 2026-01-06 Furkan Genç , Boran İsmet Macun , Sait Sarper Özaslan , Emine U. Saritas , Tolga Çukur

Incremental Decoding is an effective framework that enables the use of an offline model in a simultaneous setting without modifying the original model, making it suitable for Low-Latency Simultaneous Speech Translation. However, this…

Computation and Language · Computer Science 2024-01-12 Jiaxin Guo , Zhanglin Wu , Zongyao Li , Hengchao Shang , Daimeng Wei , Xiaoyu Chen , Zhiqiang Rao , Shaojun Li , Hao Yang

We consider the problem of independently, in a disentangled fashion, controlling the outputs of text-to-image diffusion models with color and style attributes of a user-supplied reference image. We present the first training-free,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Aishwarya Agarwal , Srikrishna Karanam , Balaji Vasan Srinivasan

We present Style Matching Score (SMS), a novel optimization method for image stylization with diffusion models. Balancing effective style transfer with content preservation is a long-standing challenge. Unlike existing efforts, our method…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Yuxin Jiang , Liming Jiang , Shuai Yang , Jia-Wei Liu , Ivor Tsang , Mike Zheng Shou

Image-to-image translation aims to learn the mapping between two visual domains. There are two main challenges for many applications: 1) the lack of aligned training pairs and 2) multiple possible outputs from a single input image. In this…

Computer Vision and Pattern Recognition · Computer Science 2018-08-03 Hsin-Ying Lee , Hung-Yu Tseng , Jia-Bin Huang , Maneesh Kumar Singh , Ming-Hsuan Yang

Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Boqiang Zhang , Hongtao Xie , Zuan Gao , Yuxin Wang

Learned Image Compression (LIC) has explored various architectures, such as Convolutional Neural Networks (CNNs) and transformers, in modeling image content distributions in order to achieve compression effectiveness. However, achieving…

Image and Video Processing · Electrical Eng. & Systems 2025-02-10 Zhuojie Wu , Heming Du , Shuyun Wang , Ming Lu , Haiyang Sun , Yandong Guo , Xin Yu

In recent years, the demand of image compression models for machine vision has increased dramatically. However, the training frameworks of image compression still focus on the vision of human, maintaining the excessive perceptual details,…

Image and Video Processing · Electrical Eng. & Systems 2025-12-24 Hyeonjin Lee , Jun-Hyuk Kim , Jong-Seok Lee

Image Super-Resolution (SR) aims to reconstruct high-resolution images from degraded low-resolution inputs. While diffusion-based SR methods offer powerful generative capabilities, their performance heavily depends on how semantic priors…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Lei Jiang , Xin Liu , Xinze Tong , Zhiliang Li , Jie Liu , Jie Tang , Gangshan Wu

Multimodal Sentiment Analysis (MSA) integrates multiple modalities to infer human sentiment, but real-world noise often leads to missing or corrupted data. However, existing feature-disentangled methods struggle to handle the internal…

Multimedia · Computer Science 2026-02-03 Xiang Li , Xiaoming Zhang , Dezhuang Miao , Xianfu Cheng , Dawei Li , Honggui Han , Zhoujun Li

LLMs have demonstrated remarkable capabilities in linguistic reasoning and are increasingly adept at vision-language tasks. The integration of image tokens into transformers has enabled direct visual input and output, advancing research…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Jonghun Kim , Sinyoung Ra , Hyunjin Park

Brain-computer interfaces (BCIs) enable direct communication between the brain and external devices, providing critical support for individuals with motor impairments. However, accurate motor imagery (MI) decoding from…

Machine Learning · Computer Science 2026-04-08 Panagiotis Andrikopoulos , Siamak Mehrkanoon

Hyperspectral image denoising faces the challenge of multi-dimensional coupling of spatially non-uniform noise and spectral correlation interference. Existing deep learning methods mostly focus on RGB images and struggle to effectively…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Haoyue Li , Di Wu

Multi-modality imaging is widely used in clinical practice and biomedical research to gain a comprehensive understanding of an imaging subject. Currently, multi-modality imaging is accomplished by post hoc fusion of independently…

Image and Video Processing · Electrical Eng. & Systems 2024-10-01 Lingting Zhu , Yizheng Chen , Lianli Liu , Lei Xing , Lequan Yu

This paper tackles the problem of disentangling the latent variables of style and content in language models. We propose a simple yet effective approach, which incorporates auxiliary multi-task and adversarial objectives, for label…

Computation and Language · Computer Science 2018-09-12 Vineet John , Lili Mou , Hareesh Bahuleyan , Olga Vechtomova

Latent Diffusion Models (LDMs) rely heavily on the compressed latent space provided by Variational Autoencoders (VAEs) for high-quality image generation. Recent studies have attempted to obtain generation-friendly VAEs by directly adopting…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 John Page , Xuesong Niu , Kai Wu , Kun Gai