Related papers: PhononBench:A Large-Scale Phonon-Based Benchmark f…
Recent image generation schemes typically capture image distribution in a pre-constructed latent space relying on a frozen image tokenizer. Though the performance of tokenizer plays an essential role to the successful generation, its…
Achieving robust performance and fairness across diverse patient populations remains a challenge in developing clinically deployable deep learning models for diagnostic imaging. Synthetic data generation has emerged as a promising strategy…
The phonon density-of-states (DOS) summarizes the lattice vibrational modes supported by a structure, and gives access to rich information about the material's stability, thermodynamic constants, and thermal transport coefficients. Here, we…
Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…
Prevalent semantic speech tokenizers, designed to capture linguistic content, are surprisingly fragile. We find they are not robust to meaning-irrelevant acoustic perturbations; even at high Signal-to-Noise Ratios (SNRs) where speech is…
High-frequency mechanical oscillators with long coherence times are essential to realizing a variety of high-fidelity quantum sensors, transducers, and memories. However, the unprecedented coherence times needed for quantum applications…
The discovery of inorganic crystal structures with targeted properties is a significant challenge in materials science. Generative models, especially state-of-the-art diffusion models, offer the promise of modeling complex data…
Efficiently generating energetically stable crystal structures has long been a challenge in material design, primarily due to the immense arrangement of atoms in a crystal lattice. To facilitate the discovery of stable material, we present…
For a very long time, computational approaches to the design of new materials have relied on an iterative process of finding a candidate material and modeling its properties. AI has played a crucial role in this regard, helping to…
The results of numerical modelling of sonic crystals with resonant array elements are reported. The investigated resonant elements include plain slotted cylinders as well as various their combinations, in particular, Russian doll or…
We study a phononic crystal interacting with an artificial atom { a superconducting quantum system { in the quantum regime. The phononic crystal is made of a long lattice of narrow metallic stripes on a quatz surface. The artificial atom in…
Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning.…
The phonon propagation dynamics in a phononic crystal waveguide, realized via a suspended one-dimensional membrane array with periodic air holes, is investigated as function of its geometry. The bandstructure of the phononic crystal can be…
The rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However,…
The ability to control phonons in solids is key for diverse quantum applications, ranging from quantum information processing to sensing. Often, phonons are sources of noise and decoherence, since they can interact with a variety of…
Uncovering the formation process that reproduces the distinct properties of compact super-Earth exoplanet systems is a major goal of planet formation theory. The most successful model argues that non-resonant systems begin as resonant…
Tool learning has generated widespread interest as a vital means of interaction between Large Language Models (LLMs) and the physical world. Current research predominantly emphasizes LLMs' capacity to utilize tools in well-structured…
Learning from noisy labels is a challenge that arises in many real-world applications where training data can contain incorrect or corrupted labels. When fine-tuning language models with noisy labels, models can easily overfit the label…
Deepfakes have become a universal and rapidly intensifying concern of generative AI across various media types such as images, audio, and videos. Among these, audio deepfakes have been of particular concern due to the ease of high-quality…
Although the tailored metal active sites and porous architectures of MOFs hold great promise for engineering challenges ranging from gas separations to catalysis, a lack of understanding of how to improve their stability limits their use in…