Related papers: A Performance Evaluation of Nomon: A Flexible Inte…
In this paper, a novel joint design of beamforming and power allocation is proposed for a multi-cell multiuser multiple-input multiple-output (MIMO) non-orthogonal multiple access (NOMA) network. In this network, base stations (BSs) adopt…
Sleep abnormalities can have severe health consequences. Automated sleep staging, i.e. labelling the sequence of sleep stages from the patient's physiological recordings, could simplify the diagnostic process. Previous work on automated…
This work proposes a novel approach for hand gesture recognition using an inexpensive, low-resolution (24 x 32) thermal sensor processed by a Spiking Neural Network (SNN) followed by Sparse Segmentation and feature-based gesture…
We propose an opinion-driven navigation framework for multi-robot traversal through a narrow corridor. Our approach leverages a multi-agent decision-making model known as the Nonlinear Opinion Dynamics (NOD) to address the narrow corridor…
Power-domain non-orthogonal multiple access (NOMA) has become a promising technology to exploit the new dimension of the power domain to enhance the spectral efficiency of wireless networks. However, most existing NOMA schemes rely on the…
Vision-and-language navigation (VLN) is a multimodal task where an agent follows natural language instructions and navigates in visual environments. Multiple setups have been proposed, and researchers apply new model architectures or…
With the ever-growing expansion of mobile technology worldwide, there is an increasing need for accommodation for those who are disabled. This project explores how machine learning and computer vision could be utilized to improve…
This research aims to find where visually impaired users find appliances hard to use and suggest guideline to solve this issue. 181 visually impaired users have been surveyed, and 12 visually impaired users have been selected based on…
Large vision-language models (VLMs) typically process hundreds or thousands of visual tokens per image or video frame, incurring quadratic attention cost and substantial redundancy. Existing token reduction methods often ignore the textual…
Non-orthogonal multiple access (NOMA) has been identified as one of the promising technologies to enhance the spectral efficiency and throughput for the fifth generation (5G) and beyond 5G cellular networks. Alternatively, Coordinated…
Distributed antenna system (DAS) has been deployed for over a decade. DAS has advantages in capacity especially for the cell edge users, in both single-cell and multi-cell environments. In this paper, non-orthogonal multiple access (NOMA)…
Crowd counting on static images is a challenging problem due to scale variations. Recently deep neural networks have been shown to be effective in this task. However, existing neural-networks-based methods often use the multi-column or…
Visible light communications (VLC) is an emerging technology with a promise of viable solution to spectrum crunch problem in conventional radio frequency (RF) bands. In this work, we consider a downlink multiuser VLC network where users…
Being able to accommodate multiple simultaneous transmissions on a single channel, non-orthogonal multiple access (NOMA) appears as an attractive solution to support massive machine type communication (mMTC) that faces a massive number of…
The explosive growth in video streaming requires video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture temporal relationships; 3D CNN-based methods can achieve good…
We propose a self-supervised shared encoder model that achieves strong results on several visual, language and multimodal benchmarks while being data, memory and run-time efficient. We make three key contributions. First, in contrast to…
Impairment of hand functions in individuals with spinal cord injury (SCI) severely disrupts activities of daily living. Recent advances have enabled rehabilitation assisted by robotic devices to augment the residual function of the muscles.…
Crowd counting usually addressed by density estimation becomes an increasingly important topic in computer vision due to its widespread applications in video surveillance, urban planning, and intelligence gathering. However, it is…
Edge deployment of large Vision-Language Models (VLMs) increasingly relies on flash-based weight offloading, where activation sparsification is used to reduce I/O overhead. However, conventional sparsification remains model-centric,…
Cross-modal alignment aims to map heterogeneous modalities into a shared latent space, as exemplified by models like CLIP, which benefit from large-scale image-text pretraining for strong recognition capabilities. However, when operating in…