Related papers: Quantifying and Correlating Rhythm Formants in Spe…
In this work we formulate a generalized theoretical model to describe the nonlinear dynamics observed in combined frequency-amplitude modulators whose characteristic parameters exhibit a nonlinear dependence on the input modulating signal.…
In this paper, we present a novel multi-modal deep neural network architecture that uses speech and text entanglement for learning phonetically sound spoken-word representations. STEPs-RL is trained in a supervised manner to predict the…
Deep learning approaches have emerged that aim to transform an audio signal so that it sounds as if it was recorded in the same room as a reference recording, with applications both in audio post-production and augmented reality. In this…
In this paper, we introduce new methods and discuss results of text-based LSTM (Long Short-Term Memory) networks for automatic music composition. The proposed network is designed to learn relationships within text documents that represent…
Morphologically rich languages accentuate two properties of distributional vector space models: 1) the difficulty of inducing accurate representations for low-frequency word forms; and 2) insensitivity to distinct lexical relations that…
Multiple scattering of wave in strong heterogeneity can cause resonance-like wave anomaly where the signal exhibits low-frequency, high intensity, and slowly propagating wave packet velocity. For example, long period event in volcanic…
In general, multi-channel source separation has utilized inter-microphone phase differences (IPDs) concatenated with magnitude information in time-frequency domain, or real and imaginary components stacked along the channel axis. However,…
In this study, we employ the atomistic wave-packet method to directly simulate coherent phonon transport and scattering dynamics in an aperiodic superlattice structure with aperiodically arranged interfaces. Our investigation reveals that…
Identification of the type of communication technology and/or modulation scheme based on detected radio signal are challenging problems encountered in a variety of applications including spectrum allocation and radio interference…
Nanomechanical resonators promise diverse applications ranging from mass spectrometry to quantum information processing, requiring long phonon lifetimes and frequency stability. Although two-level system (TLS) defects govern dissipation at…
This study characterises the radio luminosity functions (RLFs) for SFGs and AGN using statistical redshift estimation in the absence of comprehensive spectroscopic data. Sensitive radio surveys over large areas detect many sources with…
Language models (LMs) are being scaled and becoming powerful. Improving their efficiency is one of the core research topics in neural information processing systems. Tay et al. (2022) provided a comprehensive overview of efficient…
The Generation and propagation of the human voice is studied in two-dimensions using a full-body domain, using direct numerical simulation. The fluid/air in the vocal tract is modeled as a compressible and viscous fluid interacting with the…
Noise-robust automatic speech recognition (ASR) has been commonly addressed by applying speech enhancement (SE) at the waveform level before recognition. However, speech-level enhancement does not always translate into consistent…
This paper presents a simple Fourier-matching method to rigorously study resonance frequencies of a sound-hard slab with a finite number of arbitrarily shaped cylindrical holes of diameter ${\cal O}(h)$ for $h\ll1$. Outside the holes, a…
Resonant mode interactions in weakly nonlinear multi-dimensional lattices and related effects are described. We concentrate on formal description of the phenomenon and consider as examples mode interactions and evolution equations for…
On the basis of the f-deformed oscillator formalism, we propose to construct nonlinear coherent states for Hamiltonian systems having linear and quadratic terms in the the number operator by means of the two following definitions: i) as…
Large Language Models (LLMs) have emerged as powerful support tools across various natural language tasks and a range of application domains. Recent studies focus on exploring their capabilities for data annotation. This paper provides a…
This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like…
Recent advances in speech foundation models (SFMs) have enabled the direct processing of spoken language from raw audio, bypassing intermediate textual representations. This capability allows SFMs to be exposed to, and potentially respond…