中文
相关论文

相关论文: SA-OOSC: A Multimodal LLM-Distilled Semantic Commu…

200 篇论文

Semantic correspondence, the task of determining relationships between different parts of images, underpins various applications including 3D reconstruction, image-to-image translation, object tracking, and visual place recognition. Recent…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Frank Fundel , Johannes Schusterbauer , Vincent Tao Hu , Björn Ommer

Vision-based 3D Semantic Scene Completion (SSC) has received growing attention due to its potential in autonomous driving. While most existing approaches follow an ego-centric paradigm by aggregating and diffusing features over the entire…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Weihua Wang , Yubo Cui , Xiangru Lin , Zhiheng Li , Zheng Fang

Semantic communication systems for goal-oriented transmission must protect task-relevant information not only through source compression but also via physical layer mapping. Existing approaches decouple constellation design and semantic…

机器学习 · 计算机科学 2026-05-15 Albert Shaju , Christo Kurisummoottil Thomas , Mayukh Roy Chowdhury

In large part due to their implicit semantic modeling, self-supervised learning (SSL) methods have significantly increased the performance of valence recognition in speech emotion recognition (SER) systems. Yet, their large size may often…

音频与语音处理 · 电气工程与系统科学 2023-09-20 Danilo de Oliveira , Navin Raj Prabhu , Timo Gerkmann

Remote Sensing Image Change Captioning (RSICC) aims to generate natural language descriptions of surface changes between multi-temporal remote sensing images, detailing the categories, locations, and dynamics of changed objects (e.g.,…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Zhiming Wang , Mingze Wang , Sheng Xu , Yanjing Li , Baochang Zhang

Effective human-robot collaboration hinges on robust communication channels, with visual signaling playing a pivotal role due to its intuitive appeal. Yet, the creation of visually intuitive cues often demands extensive resources and…

机器人学 · 计算机科学 2024-09-24 Shubham Sonawani , Fabian Weigend , Heni Ben Amor

Out-of-context (OOC) detection is a challenging task involving identifying images and texts that are irrelevant to the context in which they are presented. Large vision-language models (LVLMs) are effective at various tasks, including image…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Fatma Shalabi , Hichem Felouat , Huy H. Nguyen , Isao Echizen

Watermarking algorithms for Large Language Models (LLMs) effectively identify machine-generated content by embedding and detecting hidden statistical features in text. However, such embedding leads to a decline in text quality, especially…

密码学与安全 · 计算机科学 2025-10-06 Yu Zhang , Shuliang Liu , Xu Yang , Xuming Hu

Task-Oriented Semantic Communication (TOSC) has been considered as a new communication paradigm to serve various samrt devices that depend on Artificial Intelligence (AI) tasks in future wireless networks. The existing TOSC frameworks rely…

信号处理 · 电气工程与系统科学 2024-07-17 Yuzhou Fu , Wenchi Cheng , Wei Zhang , Wei Zhang

Recently, very large language models (LLMs) have shown exceptional performance on several English NLP tasks with just in-context learning (ICL), but their utility in other languages is still underexplored. We investigate their effectiveness…

计算与语言 · 计算机科学 2024-06-28 Vipul Rathore , Aniruddha Deb , Ankish Chandresh , Parag Singla , Mausam

Semantic communication, when examined through the lens of joint source-channel coding (JSCC), maps source messages directly into channel input symbols, where the measure of success is defined by end-to-end distortion rather than traditional…

信息论 · 计算机科学 2024-07-09 Tze-Yang Tung , Homa Esfahanizadeh , Jinfeng Du , Harish Viswanathan

Improving the performance of semantic segmentation models using multispectral information is crucial, especially for environments with low-light and adverse conditions. Multi-modal fusion techniques pursue either the learning of…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Aniruddh Sikdar , Jayant Teotia , Suresh Sundaram

Remote sensing scene classification (RSSC) is a critical task with diverse applications in land use and resource management. While unimodal image-based approaches show promise, they often struggle with limitations such as high intra-class…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Jinjin Cai , Kexin Meng , Baijian Yang , Gang Shao

Short Text Classification (STC) is crucial for processing and understanding the brief but substantial content prevalent on contemporary digital platforms. The STC encounters difficulties in grasping the semantic and syntactic intricacies,…

计算与语言 · 计算机科学 2025-01-22 Hui Wu , Yuanben Zhang , Zhonghe Han , Yingyan Hou , Lei Wang , Siye Liu , Qihang Gong , Yunping Ge

Semantic communications have been utilized to execute numerous intelligent tasks by transmitting task-related semantic information instead of bits. In this article, we propose a semantic-aware speech-to-text transmission system for the…

音频与语音处理 · 电气工程与系统科学 2024-10-08 Zhenzi Weng , Zhijin Qin , Huiqiang Xie , Xiaoming Tao , Khaled B. Letaief

In this paper, we propose a novel semantic-aided image communication framework for supporting the compatibility with practical separation-based coding architectures. Particularly, the deep learning (DL)-based joint source-channel coding…

信息论 · 计算机科学 2025-05-01 Mingkai Xu , Yongpeng Wu , Yuxuan Shi , Xiang-Gen Xia , Merouane Debbah , Wenjun Zhang , Ping Zhang

Multimodal large language models (MLLMs) hold the potential to enhance autonomous driving by combining domain-independent world knowledge with context-specific language guidance. Their integration into autonomous driving systems shows…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Tin Stribor Sohn , Philipp Reis , Maximilian Dillitzer , Johannes Bach , Jason J. Corso , Eric Sax

Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and precision earth observation. Current multi-modal approaches mainly focus on integrating…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jinkun Dai , Yuanxin Ye , Peng Tang , Tengfeng Tang , Xianping Ma , Jing Xiao , Mi Wang

In this paper, we for the first time explore helpful multi-modal contextual knowledge to understand novel categories for open-vocabulary object detection (OVD). The multi-modal contextual knowledge stands for the joint relationship across…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Yifan Xu , Mengdan Zhang , Xiaoshan Yang , Changsheng Xu

Sign language translation (SLT) is a challenging task that involves translating sign language images into spoken language. For SLT models to perform this task successfully, they must bridge the modality gap and identify subtle variations in…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Jungeun Kim , Hyeongwoo Jeon , Jongseong Bae , Ha Young Kim