中文
相关论文

相关论文: Visual Language Model based Cross-modal Semantic C…

200 篇论文

Sign language translation (SLT) aims to convert continuous sign language videos into textual sentences. As a typical multi-modal task, there exists an inherent modality gap between sign language videos and spoken language text, which makes…

计算与语言 · 计算机科学 2023-12-27 Rui Zhao , Liang Zhang , Biao Fu , Cong Hu , Jinsong Su , Yidong Chen

This paper introduces a novel framework that synergistically combines Non-Orthogonal Multiple Access (NOMA) with Color Shift Keying (CSK) modulation to substantially boost spectral efficiency in Visible Light Communication (VLC) systems. A…

信号处理 · 电气工程与系统科学 2025-08-12 Serkan Vela , Gokce Hacioglu

We propose the Neuro-Symbolic Concept Learner (NS-CL), a model that learns visual concepts, words, and semantic parsing of sentences without explicit supervision on any of them; instead, our model learns by simply looking at images and…

计算机视觉与模式识别 · 计算机科学 2019-04-30 Jiayuan Mao , Chuang Gan , Pushmeet Kohli , Joshua B. Tenenbaum , Jiajun Wu

The rapid advancement of large language models (LLMs) has accelerated the emergence of in-context learning (ICL) as a cutting-edge approach in the natural language processing domain. Recently, ICL has been employed in visual understanding…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Dianmo Sheng , Dongdong Chen , Zhentao Tan , Qiankun Liu , Qi Chu , Jianmin Bao , Tao Gong , Bin Liu , Shengwei Xu , Nenghai Yu

The evolution of semantic communications has profoundly impacted wireless video transmission, whose applications dominate driver of modern bandwidth consumption. However, most existing schemes are predominantly optimized for simple additive…

信息论 · 计算机科学 2026-01-13 Bingyan Xie , Yongpeng Wu , Wenjun Zhang , Derrick Wing Kwan Ng , Merouane Debbah

Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, VLM in the remote…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Xingxing Weng , Chao Pang , Gui-Song Xia

Recently, numerous algorithms have been developed to tackle the problem of vision-language navigation (VLN), i.e., entailing an agent to navigate 3D environments through following linguistic instructions. However, current VLN agents simply…

计算机视觉与模式识别 · 计算机科学 2021-03-08 Hanqing Wang , Wenguan Wang , Wei Liang , Caiming Xiong , Jianbing Shen

As a paradigm shift towards pervasive intelligence, semantic communication (SemCom) has shown great potentials to improve communication efficiency and provide user-centric services by delivering task-oriented semantic meanings. However, the…

信号处理 · 电气工程与系统科学 2025-05-02 Hao Wei , Wen Wang , Wanli Ni , Wenjun Xu , Yongming Huang , Dusit Niyato , Ping Zhang

This paper proposes a novel, robust, and lightweight supervised Convolutional Neural Network (CNN)-based technique for frame identification and synchronization, designed to enhance short-link communication performance in a screen-to-camera…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Vaigai Nayaki Yokar , Hoa Le-Minh , Xicong Li , Wai Lok Woo , Luis Nero Alves , Stanislav Zvanovec , Tran The Son , Zabih Ghassemlooy

Improving the performance of semantic segmentation models using multispectral information is crucial, especially for environments with low-light and adverse conditions. Multi-modal fusion techniques pursue either the learning of…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Aniruddh Sikdar , Jayant Teotia , Suresh Sundaram

Voice conversion (VC) aims to modify the speaker's timbre while retaining speech content. Previous approaches have tokenized the outputs from self-supervised into semantic tokens, facilitating disentanglement of speech content information.…

声音 · 计算机科学 2024-09-11 Zhengyang Chen , Shuai Wang , Mingyang Zhang , Xuechen Liu , Junichi Yamagishi , Yanmin Qian

Existing research for image captioning usually represents an image using a scene graph with low-level facts (objects and relations) and fails to capture the high-level semantics. In this paper, we propose a Theme Concepts extended Image…

计算机视觉与模式识别 · 计算机科学 2021-06-22 Zhihao Fan , Zhongyu Wei , Siyuan Wang , Ruize Wang , Zejun Li , Haijun Shan , Xuanjing Huang

Semantic communications (SCs) play a central role in shaping the future of the sixth generation (6G) wireless systems, which leverage rapid advances in deep learning (DL). In this regard, end-to-end optimized DL-based joint source-channel…

信息论 · 计算机科学 2025-05-02 Mahmoud M. Salim , Mohamed S. Abdalzaher , Ali H. Muqaibel , Hussein A. Elsayed , Inkyu Lee

Vision-language navigation (VLN) is a critical domain within embedded intelligence, requiring agents to navigate 3D environments based on natural language instructions. Traditional VLN research has focused on improving environmental…

人工智能 · 计算机科学 2024-09-24 Zhiyuan Li , Yanfeng Lv , Ziqin Tu , Di Shang , Hong Qiao

It is widely agreed that open-vocabulary-based approaches outperform classical closed-set training solutions for recognizing unseen objects in images for semantic segmentation. Existing open-vocabulary approaches leverage vision-language…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Huadong Tang , Youpeng Zhao , Yan Huang , Min Xu , Jun Wang , Qiang Wu

Semantic communication (SemCom) has emerged as a promising technique for the next-generation communication systems, in which the generation at the receiver side is allowed with semantic features' recovery. However, the majority of existing…

图像与视频处理 · 电气工程与系统科学 2025-07-08 Chengyang Liang , Dong Li

The rapid development of generative Artificial Intelligence (AI) continually unveils the potential of Semantic Communication (SemCom). However, current talking-face SemCom systems still encounter challenges such as low bandwidth…

信息论 · 计算机科学 2024-11-07 Feibo Jiang , Siwei Tu , Li Dong , Cunhua Pan , Jiangzhou Wang , Xiaohu You

Semantic communication has gained significant attention from researchers as a promising technique to replace conventional communication in the next generation of communication systems, primarily due to its ability to reduce communication…

信息论 · 计算机科学 2025-02-07 Loc X. Nguyen , Ye Lin Tun , Yan Kyaw Tun , Minh N. H. Nguyen , Chaoning Zhang , Zhu Han , Choong Seon Hong

Evaluating whether vision-language models (VLMs) reason consistently across representations is challenging because modality comparisons are typically confounded by task differences and asymmetric information. We introduce SEAM, a benchmark…

人工智能 · 计算机科学 2025-08-26 Zhenwei Tang , Difan Jiao , Blair Yang , Ashton Anderson

The rapid development of artificial intelligence has driven smart health with next-generation wireless communication technologies, stimulating exciting applications in remote diagnosis and intervention. To enable a timely and effective…

机器学习 · 计算机科学 2025-08-04 Fupei Guo , Hao Zheng , Xiang Zhang , Li Chen , Yue Wang , Songyang Zhang
‹ 上一页 1 8 9 10 下一页 ›