English
Related papers

Related papers: Language-oriented Semantic Communication for Image…

200 papers

Diffusion models have been shown to be capable of generating high-quality images, suggesting that they could contain meaningful internal representations. Unfortunately, the feature maps that encode a diffusion model's internal information…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Grace Luo , Lisa Dunlap , Dong Huk Park , Aleksander Holynski , Trevor Darrell

Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a…

Traditional communication systems focus on the transmission process, and the context-dependent meaning has been ignored. The fact that 5G system has approached Shannon limit and the increasing amount of data will cause communication…

Signal Processing · Electrical Eng. & Systems 2022-02-22 Chen Dong , Haotai Liang , Xiaodong Xu , Shujun Han , Bizhu Wang , Ping Zhang

Enriching information of spectrum coverage, radiomap plays an important role in many wireless communication applications, such as resource allocation and network optimization. To enable real-time, distributed spectrum management,…

Signal Processing · Electrical Eng. & Systems 2026-01-21 Yueling Zhou , Achintha Wijesinghe , Yue Wang , Songyang Zhang , Zhipeng Cai

Diffusion models have demonstrated exceptional capabilities in generating a broad spectrum of visual content, yet their proficiency in rendering text is still limited: they often generate inaccurate characters or words that fail to blend…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Jianyi Zhang , Yufan Zhou , Jiuxiang Gu , Curtis Wigington , Tong Yu , Yiran Chen , Tong Sun , Ruiyi Zhang

Semantic image synthesis aims to generate high-quality images given semantic conditions, i.e. segmentation masks and style reference images. Existing methods widely adopt generative adversarial networks (GANs). GANs take all conditional…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Feng Liu , Xiaobin Chang

Empowered by deep learning, semantic communication marks a paradigm shift from transmitting raw data to conveying task-relevant meaning, enabling more efficient and intelligent wireless systems. In this study, we explore a deep…

Information Theory · Computer Science 2026-01-28 Chenyang Wang , Roger Olsson , Stefan Forsström , Qing He

In this paper, the problem of semantic-based efficient image transmission is studied over the Internet of Vehicles (IoV). In the considered model, a vehicle shares massive amount of visual data perceived by its visual sensors to assist…

Networking and Internet Architecture · Computer Science 2022-10-12 Qiang Pan , Haonan Tong , Jie Lv , Tao Luo , Zhilong Zhang , Changchuan Yin , Jianfeng Li

Sensing and communication are fundamental enablers of next-generation networks. While communication technologies have advanced significantly, sensing remains limited to conventional parameter estimation and is far from fully explored.…

Signal Processing · Electrical Eng. & Systems 2026-04-01 Xiaoqi Zhang , J. Andrew Zhang , Chang Liu , Weijie Yuan , Geoffrey Ye Li , Moeness G. Amin

Nowadays, the demand for image transmission over wireless networks has surged significantly. To meet the need for swift delivery of high-quality images through time-varying channels with limited bandwidth, the development of efficient…

Computational Engineering, Finance, and Science · Computer Science 2024-02-13 Mohammad Amin Jarrahi , Eirina Bourtsoulatze , Vahid Abolghasemi

Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, \textit{etc}. Additionally, these methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Hao Zhang , Lei Cao , Jiayi Ma

We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standard denoising process…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Zirui Wang , Zhizhou Sha , Zheng Ding , Yilin Wang , Zhuowen Tu

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Yunji Kim , Jiyoung Lee , Jin-Hwa Kim , Jung-Woo Ha , Jun-Yan Zhu

While Diffusion Generative Models have achieved great success on image generation tasks, how to efficiently and effectively incorporate them into speech generation especially translation tasks remains a non-trivial problem. Specifically,…

Computation and Language · Computer Science 2023-10-27 Yongxin Zhu , Zhujin Gao , Xinyuan Zhou , Zhongyi Ye , Linli Xu

Semantic communication is a promising technology to improve communication efficiency by transmitting only the semantic information of the source data. However, traditional semantic communication methods primarily focus on data…

Sound · Computer Science 2024-10-07 Jiahao Zheng , Jinke Ren , Peng Xu , Zhihao Yuan , Jie Xu , Fangxin Wang , Gui Gui , Shuguang Cui

Joint source-channel coding (JSCC) is a promising paradigm for next-generation communication systems, particularly in challenging transmission environments. In this paper, we propose a novel standard-compatible JSCC framework for the…

Information Theory · Computer Science 2025-01-07 Xue Han , Yongpeng Wu , Zhen Gao , Biqian Feng , Yuxuan Shi , Deniz Gündüz , Wenjun Zhang

This paper develops an edge-device collaborative Generative Semantic Communications (Gen SemCom) framework leveraging pre-trained Multi-modal/Vision Language Models (M/VLMs) for ultra-low-rate semantic communication via textual prompts. The…

Information Theory · Computer Science 2025-05-05 Mengmeng Ren , Li Qiao , Long Yang , Zhen Gao , Jian Chen , Mahdi Boloursaz Mashhadi , Pei Xiao , Rahim Tafazolli , Mehdi Bennis

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

Generating image variations, where a model produces variations of an input image while preserving the semantic context has gained increasing attention. Current image variation techniques involve adapting a text-to-image model to reconstruct…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Manoj Kumar , Neil Houlsby , Emiel Hoogeboom

Text-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-10-06 Hila Chefer , Oran Lang , Mor Geva , Volodymyr Polosukhin , Assaf Shocher , Michal Irani , Inbar Mosseri , Lior Wolf