English
Related papers

Related papers: MapViT: A Two-Stage ViT-Based Framework for Real-T…

200 papers

Foundation deep learning (DL) models are general models, designed to learn general, robust and adaptable representations of their target modality, enabling finetuning across a range of downstream tasks. These models are pretrained on large,…

Signal Processing · Electrical Eng. & Systems 2024-11-18 Ahmed Aboulfotouh , Ashkan Eshaghbeigi , Hatem Abou-Zeid

Radio maps (RMs), which provide location-based pathloss estimations, are fundamental to enabling proactive, environment-aware communication in 6G networks. However, existing deep learning-based methods for RM construction often model…

Networking and Internet Architecture · Computer Science 2025-11-25 Honggang Jia , Nan Cheng , Xiucheng Wang

As wireless communication technology progresses towards the sixth generation (6G), high-frequency millimeter-wave (mmWave) communication has emerged as a promising candidate for enabling vehicular networks. It offers high data rates and…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Ghazi Gharsallah , Georges Kaddoum

Aerial Vision-and-Language Navigation (AVLN) requires Unmanned Aerial Vehicle (UAV) agents to localize targets in large-scale urban environments based on linguistic instructions. While successful navigation demands both global environmental…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Xichen Ding , Jianzhe Gao , Cong Pan , Wenguan Wang , Jie Qin

As wireless communication networks rapidly evolve, spectrum resources are increasingly scarce, making effective spectrum management critically important. Radio map is a spatial representation of signal characteristics across different…

Signal Processing · Electrical Eng. & Systems 2025-12-02 Zeyao Sun , Bohao Fan , Qingyu Liu , Shuhang Zhang , Lingyang Song

Vision Transformers (ViTs) have demonstrated strong performance across a range of computer vision tasks by modeling long-range spatial interactions via self-attention. However, channel-wise mixing in ViTs remains static, relying on fixed…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Aon Safdar , Mohamed Saadeldin

In the evolving landscape of 6G networks, semantic communications are poised to revolutionize data transmission by prioritizing the transmission of semantic meaning over raw data accuracy. This paper presents a Vision Transformer…

Image and Video Processing · Electrical Eng. & Systems 2025-03-24 Muhammad Ahmed Mohsin , Muhammad Jazib , Zeeshan Alam , Muhmmad Farhan Khan , Muhammad Saad , Muhammad Ali Jamshed

Indoor pathloss prediction is a fundamental task in wireless network planning, yet it remains challenging due to environmental complexity and data scarcity. In this work, we propose a deep learning-based approach utilizing a vision…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Rafayel Mkrtchyan , Edvard Ghukasyan , Khoren Petrosyan , Hrant Khachatrian , Theofanis P. Raptis

Accurate indoor pathloss prediction is crucial for optimizing wireless communication in indoor settings, where diverse materials and complex electromagnetic interactions pose significant modeling challenges. This paper introduces…

Signal Processing · Electrical Eng. & Systems 2025-01-28 Xin Li , Ran Liu , Saihua Xu , Sirajudeen Gulam Razul , Chau Yuen

Accurate and reliable air quality forecasting is essential for protecting public health, sustainable development, pollution control, and enhanced urban planning. This letter presents a novel WaveCatBoost architecture designed to forecast…

Machine Learning · Computer Science 2025-02-18 Jintu Borah , Tanujit Chakraborty , Md. Shahrul Md. Nadzir , Mylene G. Cayetano , Shubhankar Majumdar

Radio Map Prediction (RMP), aiming at estimating coverage of radio wave, has been widely recognized as an enabling technology for improving radio spectrum efficiency. However, fast and reliable radio map prediction can be very challenging…

Signal Processing · Electrical Eng. & Systems 2021-05-18 Yu Tian , Shuai Yuan , Weisheng Chen , Naijin Liu

Next-generation radio astronomy surveys are delivering millions of resolved sources, but robust and scalable morphology analysis remains difficult across heterogeneous telescopes and imaging pipelines. We present STRADAViT, a…

Instrumentation and Methods for Astrophysics · Physics 2026-04-09 Andrea DeMarco , Ian Fenech Conti , Hayley Camilleri , Ardiana Bushi , Simone Riggi

The modeling of environmental ecosystems plays a pivotal role in the sustainable management of our planet. Accurate prediction of key environmental variables over space and time can aid in informed policy and decision-making, thus improving…

Computation and Language · Computer Science 2024-08-13 Haoran Li , Junqi Liu , Zexian Wang , Shiyuan Luo , Xiaowei Jia , Huaxiu Yao

Vision Transformers (ViTs) are essential as foundation backbones in establishing the visual comprehension capabilities of Multimodal Large Language Models (MLLMs). Although most ViTs achieve impressive performance through image-text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Weijie Yin , Dingkang Yang , Hongyuan Dong , Zijian Kang , Jiacong Wang , Xiao Liang , Chao Feng , Jiao Ran

High-level robot skills represent an increasingly popular paradigm in robot programming. However, configuring the skills' parameters for a specific task remains a manual and time-consuming endeavor. Existing approaches for learning or…

Robotics · Computer Science 2024-08-23 Claudius Kienle , Benjamin Alt , Onur Celik , Philipp Becker , Darko Katic , Rainer Jäkel , Gerhard Neumann

Speech quality and intelligibility are significantly degraded in noisy environments. This paper presents a novel transformer-based learning framework to address the single-channel noise suppression problem for real-time applications.…

Sound · Computer Science 2025-11-18 Behnaz Bahmei , Siamak Arzanpour , Elina Birmingham

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

Timely information delivery in low-altitude networks is critical for many time-sensitive applications, such as unmanned aerial vehicle (UAV) navigation, inspection, and surveillance. The key challenge lies in balancing three competing…

Systems and Control · Electrical Eng. & Systems 2026-04-30 Bowen Li , Jiping Luo , Themistoklis Charalambous , Nikolaos Pappas

In recent years, a variety of ML architectures and techniques have seen success in producing skillful medium range weather forecasts. In particular, Vision Transformer (ViT)-based models (e.g. Pangu-Weather, FuXi) have shown strong…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Vivek Ramavajjala

General-purpose pre-trained models ("foundation models") have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that are significantly smaller than those required for learning…

Robotics · Computer Science 2023-10-25 Dhruv Shah , Ajay Sridhar , Nitish Dashora , Kyle Stachowicz , Kevin Black , Noriaki Hirose , Sergey Levine
‹ Prev 1 2 3 10 Next ›