Related papers: Generating Localized Audible Zones Using a Single-…
Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data…
We prove stability and exponential convergence of the Perfectly Matched Layer (PML) method for acoustic scattering on manifolds with axial analytic quasicylindrical ends. These manifolds model long-range geometric perturbations (e.g.…
It is known that any {\em real coordinate transformation} (RCT) to compress waves in an unbounded domain into a bounded domain results in infinite oscillations that cannot be resolved by any grid-based method. In this paper, we intend to…
Integrated sensing and communication (ISAC) at terahertz (THz) frequencies holds significant promise for unifying ultra-high-speed wireless connectivity with fine-grained environmental awareness. Realistic and interpretable channel modeling…
In this paper, we consider the Spatial Modulation (SM) system in a frequency selective channel under single carrier (SC) communication scenario and propose zero-padding instead of cyclic prefix considered in the existing literature. We show…
We develop a low-complexity coding scheme to achieve covert communications over binary-input discrete memoryless channels (BI-DMCs). We circumvent the impossibility of covert communication with linear codes by introducing non-linearity…
In this paper, two modulation schemes based on complementary sequences (CSs) are proposed for uplink control channels in unlicensed bands. These schemes address high peak-to-average-power ratio (PAPR) under non-contiguous resource…
Speech separation has been shown effective for multi-talker speech recognition. Under the ad hoc microphone array setup where the array consists of spatially distributed asynchronous microphones, additional challenges must be overcome as…
A binaural rendering framework for personal sound zones (PSZs) is proposed to enable multiple head-tracked listeners to receive fully independent stereo audio programs. Current PSZ systems typically rely on monophonic rendering and…
Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task.…
Aiming for the sixth generation (6G) wireless communications, distributed massive multiple-input multiple-output (MIMO) systems hold significant potential for spatial multiplexing. In order to evaluate the ability of a distributed massive…
We present a system for the Zero Resource Speech Challenge 2021, which combines a Contrastive Predictive Coding (CPC) with deep cluster. In deep cluster, we first prepare pseudo-labels obtained by clustering the outputs of a CPC network…
The increasing presence of large-scale distributed systems highlights the need for scalable control strategies where only local communication is required. Moreover, in safety-critical systems it is imperative that such control strategies…
Large Language Models (LLMs) excel at generating fluent text but struggle to enforce external constraints because they generate tokens sequentially without explicit control mechanisms. GenCP addresses this limitation by combining LLM…
Space-division multiplexing is a promising technology in optical fibre communication to improve the transmission capacity of a single optical fibre. However, the number of channels that can be multiplexed is limited by the crosstalks…
Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly…
Multichannel active noise control (ANC) systems are designed to create a large zone of quietness (ZoQ) around the error microphones, however, the placement of these microphones often presents challenges due to physical limitations. Virtual…
Joint communication and localization~(JCL) is envisioned to be a key feature in future millimeter-wave~(mmWave) wireless networks for context-aware applications. A map-based channel model considering both site-specific radio environment and…
Large language models (LLMs) have spurred development in multiple industries. However, the growing number of their parameters brings substantial storage and computing burdens, making it essential to explore model compression techniques for…
This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current speaker models and…