English
Related papers

Related papers: Context Video Semantic Transmission with Variable …

200 papers

The success of large-scale language models has established tokens as compact and meaningful units for natural-language representation, which motivates token communication over wireless channels, where tokens are considered fundamental units…

Signal Processing · Electrical Eng. & Systems 2026-01-27 Junyong Shin , Joohyuk Park , Jihong Park , Jinho Choi , Yo-Seb Jeon

Subword tokenization is a common method for vocabulary building in Neural Machine Translation (NMT) models. However, increasingly complex tasks have revealed its disadvantages. First, a vocabulary cannot be modified once it is learned,…

Computation and Language · Computer Science 2024-08-13 Langlin Huang , Yang Feng

In this paper, we propose a novel cross-layer ap-proach for video delivery over Wireless Video Sensor Networks (WVSN)s. We adopt an energy efficient and adaptive video compression scheme dedicated to the WVSNs, based on the H.264/AVC video…

Networking and Internet Architecture · Computer Science 2015-01-30 Othmane Alaoui-Fdili , Patrick CORLAY , Youssef Fakhri , François-Xavier Coudoux , Driss Aboutajdine

Semantic communication has emerged as a promising paradigm for next-generation networks, yet several fundamental challenges remain unresolved. Building on the probabilistic model of semantic communication and leveraging the concept of…

Information Theory · Computer Science 2026-02-27 Javad Gholipour , Rafael F. Schaefer , Gerhard P. Fettweis

Massive multiple-input multiple-output (MIMO) stands as a key technology for advancing performance metrics such as data rate, reliability, and spectrum efficiency in the fifth generation (5G) and beyond of wireless networks. However, its…

Signal Processing · Electrical Eng. & Systems 2025-05-19 Sajjad Nassirpour , Toan-Van Nguyen , Duy H. N. Nguyen

Compressive subspace learning (CSL) with the exploitation of space diversity has found a potential performance improvement for wideband spectrum sensing (WBSS). However, previous works mainly focus on either exploiting antenna…

Information Theory · Computer Science 2020-06-09 Tierui Gong , Zhijia Yang , Meng Zheng , Zhifeng Liu , Gengshan Wang

Video-grounded dialogue understanding is a challenging problem that requires machine to perceive, parse and reason over situated semantics extracted from weakly aligned video and dialogues. Most existing benchmarks treat both modalities the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Yuxuan Wang , Zilong Zheng , Xueliang Zhao , Jinpeng Li , Yueqian Wang , Dongyan Zhao

Variational Neural Machine Translation (VNMT) is an attractive framework for modeling the generation of target translations, conditioned not only on the source sentence but also on some latent random variables. The latent variable modeling…

Computation and Language · Computer Science 2020-05-29 Hendra Setiawan , Matthias Sperber , Udhay Nallasamy , Matthias Paulik

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

Sound · Computer Science 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

Large vision models like the Segment Anything Model (SAM) exhibit significant limitations when applied to downstream tasks in the wild. Consequently, reference segmentation, which leverages reference images and their corresponding masks to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Haoran Wang , Zekun Li , Jian Zhang , Lei Qi , Yinghuan Shi

Though significant progress has been made for speaker-dependent Video-to-Speech (VTS) synthesis, little attention is devoted to multi-speaker VTS that can map silent video to speech, while allowing flexible control of speaker identity, all…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-21 Disong Wang , Shan Yang , Dan Su , Xunying Liu , Dong Yu , Helen Meng

Deep learning based semantic communication (DeepSC) system has emerged as a promising paradigm for efficient wireless transmission. However, existing image DeepSC methods, frequently encounter challenges in balancing rate-distortion…

Image and Video Processing · Electrical Eng. & Systems 2025-12-08 Yinhuan Huang , Zhijin Qin

The recent success of interleaved Large Multimodal Models (LMMs) in few-shot learning suggests that in-context learning (ICL) with many examples can be promising for learning new tasks. However, this many-shot multimodal ICL setting has one…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Brandon Huang , Chancharik Mitra , Assaf Arbelle , Leonid Karlinsky , Trevor Darrell , Roei Herzig

In this paper, we develop a deep learning based semantic communication system for speech transmission, named DeepSC-ST. We take the speech recognition and speech synthesis as the transmission tasks of the communication system, respectively.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-03 Zhenzi Weng , Zhijin Qin , Xiaoming Tao , Chengkang Pan , Guangyi Liu , Geoffrey Ye Li

Context-aware methods have achieved remarkable advancements in supervised scene text recognition by leveraging semantic priors from words. Considering the heterogeneity of text and background in STR, we propose that such contextual priors…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Tiancheng Lin , Jinglei Zhang , Yi Xu , Kai Chen , Rui Zhang , Chang-Wen Chen

This paper proposes a network, referred to as MVSTR, for Multi-View Stereo (MVS). It is built upon Transformer and is capable of extracting dense features with global context and 3D consistency, which are crucial to achieving reliable…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Jie Zhu , Bo Peng , Wanqing Li , Haifeng Shen , Zhe Zhang , Jianjun Lei

Vision-Language models (VLMs) have excelled in the image-domain -- especially in zero-shot settings -- thanks to the availability of vast pretraining data (i.e., paired image-text samples). However for videos, such paired data is not as…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Kumara Kahatapitiya , Anurag Arnab , Arsha Nagrani , Michael S. Ryoo

Semantic communication has undergone considerable evolution due to the recent rapid development of artificial intelligence (AI), significantly enhancing both communication robustness and efficiency. Despite these advancements, most current…

Image and Video Processing · Electrical Eng. & Systems 2024-05-24 Jiarun Ding , Peiwen Jiang , Chao-Kai Wen , Shi Jin

The growing demand for efficient semantic communication systems capable of managing diverse tasks and adapting to fluctuating channel conditions has driven the development of robust, resource-efficient frameworks. This article introduces a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Xiang Chen , Shuying Gan , Chenyuan Feng , Xijun Wang , Tony Q. S. Quek

Collaborative perception, an emerging paradigm in autonomous driving, has been introduced to mitigate the limitations of single-vehicle systems, such as limited sensor range and occlusion. To improve the robustness of inter-vehicle data…

Signal Processing · Electrical Eng. & Systems 2025-11-26 Mingyi Lu , Guowei Liu , Le Liang , Chongtao Guo , Hao Ye , Shi Jin
‹ Prev 1 8 9 10 Next ›