English
Related papers

Related papers: Object-Attribute-Relation Model Driven Adaptive Hi…

200 papers

Large Language Models (LLMs) have made significant strides in text generation and comprehension, with recent advancements extending into multimodal LLMs that integrate visual and audio inputs. However, these models continue to struggle with…

Computation and Language · Computer Science 2024-10-17 Arushi Goel , Karan Sapra , Matthieu Le , Rafael Valle , Andrew Tao , Bryan Catanzaro

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Moayed Haji-Ali , Willi Menapace , Aliaksandr Siarohin , Ivan Skorokhodov , Alper Canberk , Kwot Sin Lee , Vicente Ordonez , Sergey Tulyakov

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Eitan Shaar , Ariel Shaulov , Gal Chechik , Lior Wolf

The integration of artificial intelligence (AI) with the Internet of Things (IoT) enables task-oriented communication for multi-edge cooperative inference system, where edge devices transmit extracted features of local sensory data to an…

Signal Processing · Electrical Eng. & Systems 2025-10-28 Dongwon Kim , Jiwan Seo , Joonhyuk Kang

The rapid development of artificial intelligence has significantly advanced semantic communications, particularly in wireless image transmission. However, most existing approaches struggle to precisely distinguish and prioritize image…

Image and Video Processing · Electrical Eng. & Systems 2024-10-24 Jiarun Ding , Peiwen Jiang , Chao-Kai Wen , Shi Jin

Foundation models in language and vision benefit from a unified discrete token interface that converts raw inputs into sequences for scalable pre-training and inference. For graphs, an effective tokenizer should yield reusable discrete…

Information Retrieval · Computer Science 2026-05-28 Yang Xiang , Li Fan , Chenke Yin , Lutz Oettershagen , Chengtao Ji

Semantic communication (SemCom) significantly reduces redundant data and improves transmission efficiency by extracting the latent features of information. However, most of the conventional deep learning-based SemCom systems focus on analog…

Signal Processing · Electrical Eng. & Systems 2025-12-30 Ming Lyu , Hao Chen , Dan Wang , Chen Qiu , Guangyin Feng , Nan Ma , Xiaodong Xu

Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information struggle with complex scenes and generating audio tailored to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yingshan Liang , Keyu Fan , Zhicheng Du , Yiran Wang , Qingyang Shi , Xinyu Zhang , Jiasheng Lu , Peiwu Qin

Multimodal affective computing, learning to recognize and interpret human affects and subjective information from multiple data sources, is still challenging because: (i) it is hard to extract informative features to represent human affects…

Computation and Language · Computer Science 2018-05-23 Yue Gu , Kangning Yang , Shiyu Fu , Shuhong Chen , Xinyu Li , Ivan Marsic

Mobile Edge Computing (MEC) offers low-latency and high-bandwidth support for Internet-of-Vehicles (IoV) applications. However, due to high vehicle mobility and finite communication coverage of base stations, it is hard to maintain…

Networking and Internet Architecture · Computer Science 2025-03-19 Zheyi Chen , Sijin Huang , Geyong Min , Zhaolong Ning , Jie Li , Yan Zhang

Hypothesis. Artificial general intelligence is, at its core, a compression problem. Effective compression demands resonance: deep learning scales best when its architecture aligns with the fundamental structure of the data. These are the…

Convolutional neural networks (CNNs) have demonstrated strong performance in visual recognition tasks, but their inherent reliance on regular grid structures limits their capacity to model complex topological relationships and non-local…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Feiyue Zhao , Zhichao Zhang

Speech is one of the most effective ways of communication among humans. Even though audio is the most common way of transmitting speech, very important information can be found in other modalities, such as vision. Vision is particularly…

Computation and Language · Computer Science 2016-11-22 Ramon Sanabria , Florian Metze , Fernando De La Torre

We present DeepWiVe, the first-ever end-to-end joint source-channel coding (JSCC) video transmission scheme that leverages the power of deep neural networks (DNNs) to directly map video signals to channel symbols, combining video…

Image and Video Processing · Electrical Eng. & Systems 2021-11-29 Tze-Yang Tung , Deniz Gündüz

Multimodal Large Language Models have achieved significant success in offline video understanding, yet their application to streaming videos is severely limited by the linear explosion of visual tokens, which often leads to Out-of-Memory…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Chao Wang , Xudong Tan , Jianjian Cao , Kangcong Li , Tao Chen

Recent deep-learning-based video compression methods brought coding gains over conventional codecs such as AVC and HEVC. However, learning-based codecs generally require considerable computation time and model complexity. In this paper, we…

Image and Video Processing · Electrical Eng. & Systems 2023-03-22 Hochang Rhee , Seyun Kim , Nam Ik Cho

In an adaptive bitrate streaming application, the efficiency of video compression and the encoded video quality depend on both the video codec and the quality metric used to perform encoding optimization. The development of such a quality…

Image and Video Processing · Electrical Eng. & Systems 2024-04-11 Yixu Chen , Zaixi Shang , Hai Wei , Yongjun Wu , Sriram Sethuraman

Collaborative perception allows connected vehicles to exchange sensor information and overcome each vehicle's blind spots. Yet transmitting raw point clouds or full feature maps overwhelms Vehicle-to-Vehicle (V2V) communications, causing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Melih Yazgan , Allen Xavier Arasan , J. Marius Zöllner

Semantic communication (SemCom) aims to achieve high fidelity information delivery under low communication consumption by only guaranteeing semantic accuracy. Nevertheless, semantic communication still suffers from unexpected channel…

Systems and Control · Electrical Eng. & Systems 2024-03-26 Fei Ni , Rongpeng Li , Zhifeng Zhao , Honggang Zhang

We present a new recurrent neural network topology to enhance state-of-the-art machine learning systems by incorporating a broader context. Our approach overcomes recent limitations with extended narratives through a multi-layered…

Computation and Language · Computer Science 2018-08-07 Patrick Huber , Jan Niehues , Alex Waibel