Related papers: Local distillation from Reed Muller codes unfoldin…
Qudits offer the potential for low-overhead magic state distillation, although previous results for asymptotically good codes have required qudit dimension $q\gg 100$ or code length $\mathcal{N}\gg 100$. These parameters far exceed…
Knowledge distillation (KD) is a key technique for compressing large language models into smaller ones while preserving performance. Despite the recent traction of KD research, its effectiveness for smaller language models (LMs) and the…
We investigate 1D and 2D cluster states under local decoherence to assess the robustness of their mixed-state subsystem symmetry-protected topological (SSPT) order. By exactly computing fidelity correlators via dimensional reduction of…
Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation models for specific medical scenarios, it remains the inevitable challenge due to the gap…
We develop a topological theory for fault-tolerant quantum computation in quantum low-density parity-check (qLDPC) codes. We show that there exist hidden simplicial or CW complex structures encoding the topological data for all qLDPC and…
We classify, up to local unitary equivalence, local unitary stabilizer Lie algebras for symmetric mixed states into six classes. These include the stabilizer types of the Werner states, the GHZ state and its generalizations, and Dicke…
A two-dimensional (2D) mathematical model of quadratically distorted (QD) grating is established with the principles of Fraunhofer diffraction and Fourier optics. Discrete sampling and bisection algorithm are applied for finding numerical…
Fully localised patterns involving cellular hexagons or squares have been found experimentally and numerically in various continuum models. However, there is currently no mathematical theory for the emergence of these localised cellular…
Previous knowledge distillation (KD) methods for object detection mostly focus on feature imitation instead of mimicking the prediction logits due to its inefficiency in distilling the localization information. In this paper, we investigate…
Knowledge distillation (KD) transfers capabilities from large language models (LLMs) to smaller students, yet it can fail unpredictably and also underpins model leakage risks. Our analysis revealed several distillation traps: tail noise,…
From an appropriate parameterization of the three-dimensional (3D) coherency matrix R, that characterizes the second-order, classical states of polarization, the coherency matrices are classified and interpreted in terms of incoherent…
Preparation of high-fidelity logical magic states has remained as a necessary but daunting step towards building a large-scale fault-tolerant quantum computer. One approach is to fault-tolerantly prepare a magic state in one code and then…
We propose {\it measurement-producing hierarchy} emerging among correlated states by sequential subsystem projective measurements. We start from symmetry-protected-topological (SPT) cluster states with a large symmetry and apply sequential…
For a broad family of discriminative models that includes autoregressive language models, identifiability results imply that if two models induce the same conditional distributions, then their internal representations agree up to an…
Medical foundation models pre-trained on large-scale datasets have demonstrated powerful versatile capabilities for various tasks. However, due to the gap between pre-training tasks (or modalities) and downstream tasks (or modalities), the…
Learning generative models directly from corrupted observations is a long standing challenge across natural and scientific domains. We introduce Restoration Score Distillation (RSD), a unified framework for learning high fidelity, one step…
In this paper, we investigate how model distillation impacts the development of reasoning features in large language models (LLMs). To explore this, we train a crosscoder on Qwen-series models and their fine-tuned variants. Our results…
LiDAR odometry and localization are two widely used and fundamental applications in robotic and autonomous driving systems. Although state-of-the-art (SOTA) systems achieve high accuracy on clean point clouds, their robustness to corrupted…
We propose a general framework for studying two-dimensional (2D) topologically ordered states subject to local correlated errors and show that the resulting mixed-state can display intrinsically mixed-state topological order (imTO) --…
Knowledge distillation compresses a larger neural model (teacher) into smaller, faster student models by training them to match teacher outputs. However, the internal computational transformations that occur during this process remain…