Related papers: Optimizing phonon space in the phonon-coupling mod…
Sound source localization (SSL) determines the position of sound sources using multi-channel audio data. It is commonly used to improve speech enhancement and separation. Extracting spatial features is crucial for SSL, especially in…
Digital quantum simulation of electron-phonon systems requires truncating infinite phonon levels into $N$ basis states and then encoding them with qubit computational basis. Unary encoding and the more compact binary/Gray encoding are the…
Inference-time alignment enhances the performance of large language models without requiring additional training or fine-tuning but presents challenges due to balancing computational efficiency with high-quality output. Best-of-N (BoN)…
Large language models (LLMs) improve reasoning accuracy when generating multiple candidate solutions at test time, but standard methods like Best-of-N (BoN) incur high computational cost by fully generating all branches. Self-Truncation…
Weak localization has a strong influence on both the normal and superconducting properties of metals. In particular, since weak localization leads to the decoupling of electrons and phonons, the temperature dependence of resistance (i.e.,…
Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cost remains high when audio inputs are represented as long prefix-token sequences. These…
Exact-binary encoding compiles a discrete cost function network (CFN) into a higher-order unconstrained binary optimization (HUBO) problem whose maximum monomial degree grows with the cardinalities of the underlying CFN variables. Given…
Audio tagging is an important task of mapping audio samples to their corresponding categories. Recently endeavours that exploit transformer models in this field have achieved great success. However, the quadratic self-attention cost limits…
The t-t'-t''-J model of electrons interacting with three phonon modes (breathing, apical breathing, and buckling) is considered. The wave-vector dependence of the matrix elements of the electron-phonon interaction leads to opposite…
Large language models (LLMs) and multimodal LLMs (MLL-Ms) excel at chain-of-thought reasoning but face distribution shift at test-time and a lack of verifiable supervision. Recent test-time reinforcement learning (TTRL) methods derive…
This paper deals with the solution of the spherically symmetric time-dependent Hartree-Fock approximation applied in the case of nuclear giant monopole resonances. The problem is spatially unbounded as the resonance state is in the…
The problems related to the existence of the spurious dipole mode (SDM) in the self-consistent nuclear-structure models are considered. A method is formulated that allows to eliminate coupling of the SDM with the physical modes in the…
The joint design of analog beamforming and power allocation is investigated for a single radio-frequency chain multiuser time-division multiple access system under a max-min signal-to-noise ratio (SNR) criterion. A hardware-efficient…
The self-consistent random phase approximation (RPA) based on a correlated realistic nucleon-nucleon interaction is used to evaluate correlation energies in closed-shell nuclei beyond the Hartree-Fock level. The relevance of contributions…
The performance of deep learning-based multi-channel speech enhancement methods often deteriorates when the geometric parameters of the microphone array change. Traditional approaches to mitigate this issue typically involve training on…
Magic-angle twisted bilayer graphene (TBG) has attracted significant interest recently due to the discoveries of diverse correlated and topological states in this system. Despite the extensive research on the electron-electron interaction…
We formulate a first-principle scheme for structural optimization at finite temperature ($T$) based on the self-consistent phonon (SCP) theory, which accurately takes into account the effect of strong phonon anharmonicity. The…
We calculate the phonon-dispersion relations of several two-dimensional materials and diamond using the density-functional based tight-binding approach (DFTB). Our goal is to verify if this numerically efficient method provides sufficiently…
Reinforcement learning (RL) is a critical component of large language model (LLM) post-training. However, on-policy algorithms used for post-training are not naturally robust to a diversified content of experience replay buffers, which…
Speech enhancement (SE) is critical for improving speech intelligibility and quality in real-world environments, particularly for cochlear implant (CI) users who experience severe degradations in speech understanding under noisy and…