中文
相关论文

相关论文: Visual Language Model based Cross-modal Semantic C…

200 篇论文

Existing Multimodal Large Language Models (MLLMs) suffer from increased inference costs due to the additional vision tokens introduced by image inputs. In this work, we propose Visual Consistency Learning (ViCO), a novel training algorithm…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Long Cui , Weiyun Wang , Jie Shao , Zichen Wen , Gen Luo , Linfeng Zhang , Yanting Zhang , Yu Qiao , Wenhai Wang

Currently, large language models (LLMs) predominantly focus on the text modality. To enable more natural human-AI interaction, speech LLMs are emerging, but building effective end-to-end speech LLMs remains challenging due to limited data…

计算与语言 · 计算机科学 2026-04-14 Yan Zhou , Qingkai Fang , Yun Hong , Yang Feng

This paper proposes a novel knowledge-Base (KB) assisted semantic communication framework for image transmission. At the receiver, a Facebook AI Similarity Search (FAISS) based vector database is constructed by extracting semantic…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Chongyang Li , Yanmei He , Tianqian Zhang , Mingjian He , Shouyin Liu

In the new paradigm of semantic communication (SC), the focus is on delivering meanings behind bits by extracting semantic information from raw data. Recent advances in data-to-text models facilitate language-oriented SC, particularly for…

计算机视觉与模式识别 · 计算机科学 2024-05-17 Giordano Cicchetti , Eleonora Grassucci , Jihong Park , Jinho Choi , Sergio Barbarossa , Danilo Comminiello

Concept Bottleneck Models (CBMs) provide interpretable prediction by introducing an intermediate Concept Bottleneck Layer (CBL), which encodes human-understandable concepts to explain models' decision. Recent works proposed to utilize Large…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Divyansh Srivastava , Ge Yan , Tsui-Wei Weng

Modern communications are usually designed to pursue a higher bit-level precision and fewer bits while transmitting a message. This article rethinks these two major features and introduces the concept and advantage of semantics that…

信号处理 · 电气工程与系统科学 2022-06-09 Kun Lu , Qingyang Zhou , Rongpeng Li , Zhifeng Zhao , Xianfu Chen , Jianjun Wu , Honggang Zhang

Recently, learning-based semantic communication (SemCom) has emerged as a promising approach in the upcoming 6G network and researchers have made remarkable efforts in this field. However, existing works have yet to fully explore the…

信号处理 · 电气工程与系统科学 2024-04-01 Shunpu Tang , Qianqian Yang , Deniz Gündüz , Zhaoyang Zhang

Semantic communications have shown promising advancements by optimizing source and channel coding jointly. However, the dynamics of these systems remain understudied, limiting research and performance gains. Inspired by the robustness of…

信号处理 · 电气工程与系统科学 2026-04-29 Hanju Yoo , Linglong Dai , Songkuk Kim , Chan-Byoung Chae

Encoder, decoder and knowledge base are three major components for semantic communication. Recent advances have achieved significant progress in the encoder-decoder design. However, there remains a considerable gap in the construction and…

网络与互联网体系结构 · 计算机科学 2025-03-18 Zhiyuan Xi , Kun Zhu , Yuanyuan Xu , Tong Zhang

Semantic communications focus on the transmission of semantic features. In this letter, we consider a task-oriented multi-user semantic communication system for multimodal data transmission. Particularly, partial users transmit images while…

信号处理 · 电气工程与系统科学 2021-12-15 Huiqiang Xie , Zhijin Qin , Geoffrey Ye Li

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

多媒体 · 计算机科学 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

As the real propagation environment becomes in creasingly complex and dynamic, millimeter wave beam prediction faces huge challenges. However, the powerful cross modal representation capability of vision-language model (VLM) provides a…

信号处理 · 电气工程与系统科学 2025-08-18 Ji Wang , Bin Tang , Jian Xiao , Qimei Cui , Xingwang Li , Tony Q. S. Quek

Thanks to the rise of deep learning and the availability of large-scale audio-visual databases, recent advances have been achieved in Visual Speech Recognition (VSR). Similar to other speech processing tasks, these end-to-end VSR systems…

计算机视觉与模式识别 · 计算机科学 2024-02-21 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Multimodal semantic communication has great potential to enhance downstream task performance by integrating complementary information across modalities. This paper introduces ProMSC-MIS, a novel Prompt-based Multimodal Semantic…

多媒体 · 计算机科学 2025-08-28 Haoshuo Zhang , Yufei Bo , Meixia Tao

Texts in scene images convey critical information for scene understanding and reasoning. The abilities of reading and reasoning matter for the model in the text-based visual question answering (TextVQA) process. However, current TextVQA…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Chengyang Fang , Gangyan Zeng , Yu Zhou , Daiqing Wu , Can Ma , Dayong Hu , Weiping Wang

Although semantic communication (SC) has shown its potential in efficiently transmitting multimodal data such as texts, speeches and images, SC for videos has focused primarily on pixel-level reconstruction. However, these SC systems may be…

计算机视觉与模式识别 · 计算机科学 2025-02-12 Jiangyuan Guo , Wei Chen , Yuxuan Sun , Jialong Xu , Bo Ai

Recent advances in semantic communication (SC) have introduced neural network (NN)-based transceivers that convey semantic representation (SR) of signals such as images. However, these NNs are trained over diverse image distributions and…

信号处理 · 电气工程与系统科学 2025-07-24 Hyelin Nam , Jihong Park , Jinho Choi , Seong-Lyun Kim

Concept bottleneck models (CBMs) improve neural network interpretability by introducing an intermediate layer that maps human-understandable concepts to predictions. Recent work has explored the use of vision-language models (VLMs) to…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Xingbo Du , Qiantong Dou , Lei Fan , Rui Zhang

Semantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the…

音频与语音处理 · 电气工程与系统科学 2024-10-08 Zhenzi Weng , Zhijin Qin , Huiqiang Xie , Xiaoming Tao , Khaled B. Letaief

Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reasoning tasks that require multimodal or multilingual…

机器学习 · 计算机科学 2025-07-09 Wenyi Wu , Zixuan Song , Kun Zhou , Yifei Shao , Zhiting Hu , Biwei Huang