English
Related papers

Related papers: MarginGate: Sparse Margin-Triggered Verification f…

200 papers

Speculative decoding (SD) is a widely used approach for accelerating decode-heavy LLM inference workloads. While online inference workloads are highly dynamic, existing SD systems are rigid and take a coarse-grained approach to SD…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-23 Wenyan Chen , Chengzhi Lu , Yanying Lin , Dmitrii Ustiugov

We construct a gate and time-independent noise model that results in the output of a logical randomized benchmarking protocol oscillating rather than decaying exponentially. To illustrate our idea, we first construct an example in standard…

Quantum Physics · Physics 2022-12-13 Athena Ceasura , Pavithran Iyer , Joel J. Wallman , Hakop Pashayan

Large language models can generate fluent answers that are unfaithful to the provided context, while many safeguards rely on external verification or a separate judge after generation. We introduce \emph{internal flow signatures} that audit…

Machine Learning · Computer Science 2026-02-03 Sungheon Jeong , Sanggeon Yun , Ryozo Masukawa , Wenjun Haung , Hanning Chen , Mohsen Imani

Financial institutions deploy Large Language Models (LLMs) for reconciliations, regulatory reporting, and client communications, but nondeterministic outputs (output drift) undermine auditability and trust. We quantify drift across five…

Machine Learning · Computer Science 2025-11-12 Raffi Khatchadourian , Rolando Franco

Selecting the appropriate model at inference time -- the routing problem -- requires jointly optimizing output quality, cost, latency, and governance constraints. Existing approaches delegate this decision to LLM-based classifiers or…

Networking and Internet Architecture · Computer Science 2026-04-06 Warren Johnson , Charles Lee

Although qubit coherence times and gate fidelities are continuously improving, logical encoding is essential to achieve fault tolerance in quantum computing. In most encoding schemes, correcting or tracking errors throughout the computation…

We introduce the Byte Latent Transformer (BLT), a new byte-level LLM architecture that, for the first time, matches tokenization-based LLM performance at scale with significant improvements in inference efficiency and robustness. BLT…

We introduce XFP, a dynamic weight quantizer for LLM inference that inverts the conventional workflow: the operator specifies reconstruction quality floors on per-channel cosine similarity (one strict floor for attention and shared experts,…

Machine Learning · Computer Science 2026-05-15 Thomas Witt

We present measurements of single-qubit gate errors for a superconducting qubit. Results from quantum process tomography and randomized benchmarking are compared with gate errors obtained from a double pi pulse experiment. Randomized…

Mesoscale and Nanoscale Physics · Physics 2009-03-08 J. M. Chow , J. M. Gambetta , L. Tornberg , Jens Koch , Lev S. Bishop , A. A. Houck , B. R. Johnson , L. Frunzio , S. M. Girvin , R. J. Schoelkopf

We estimate and analyze the error rates and the resource overheads of the repetition cat qubit approach to universal and fault-tolerant quantum computation. The cat qubits stabilized by two-photon dissipation exhibit an extremely biased…

Quantum Physics · Physics 2021-04-21 Jérémie Guillaud , Mazyar Mirrahimi

Recent research on the 1-bit Large Language Models (LLMs), such as BitNet b1.58, presents a promising direction for reducing the inference cost of LLMs while maintaining their performance. In this work, we introduce BitNet a4.8, enabling…

Computation and Language · Computer Science 2024-11-08 Hongyu Wang , Shuming Ma , Furu Wei

Fault-tolerant logical operations for qubits encoded by CSS codes are discussed, with emphasis on methods that apply to codes of high rate, encoding k qubits per block with k>1. It is shown that the logical qubits within a given block can…

Quantum Physics · Physics 2013-05-29 Andrew M. Steane , Ben Ibinson

While instruction fine-tuned LLMs are effective text generators, sensitivity to prompt construction makes performance unstable and sub-optimal in practice. Relying on a single "best" prompt cannot capture all differing approaches to a…

Computation and Language · Computer Science 2024-10-07 David Heineman , Yao Dou , Wei Xu

Speculative decoding has emerged as a promising technique for large language model (LLM) inference by accelerating autoregressive decoding via draft-then-verify. This paper studies a new edge scenario with multi-user inference, where draft…

Information Theory · Computer Science 2026-04-24 Yaodan Xu , Sheng Zhou , Zhisheng Niu

Quantization of Large Language Models (LLMs) has recently gained popularity, particularly for on-device settings with limited hardware resources. While efficient, quantization inevitably degrades model quality, especially in aggressive…

Machine Learning · Computer Science 2025-06-25 Yeonhong Park , Jake Hyun , Hojoon Kim , Jae W. Lee

In state-of-the-art quantum computing platforms, including superconducting qubits and trapped ions, imperfections in the 2-qubit entangling gates are the dominant contributions of error to system-wide performance. Recently, a novel 2-qubit…

Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary filters often miss irreversible mid-trajectory tool calls, while post-hoc LLM judges…

Computation and Language · Computer Science 2026-05-27 Haoxuan Jia , Yang Liu , Bin Chong , Yingguang Yang , Yancheng Chen , Jiayu Liang , Qian Li , Hanning Lu , Kefu Xu , Hao Zheng , Chongyang Zhang , Hao Peng , Philip S. Yu

Medical Vision Language Models (VLMs) can change their answers when clinicians rephrase the same question, a failure mode that threatens deployment safety. We introduce PSF-Med, a benchmark of 26,850 chest X-ray questions paired with 92,856…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Binesh Sadanandan , Vahid Behzadan

Microwave-driven logic is a promising alternative to laser control in scaling trapped-ion based quantum processors. However, such electronic gates have yet to match the speed offered by their laser-driven counterparts. Here, we implement…

Large Language Models (LLMs) are increasingly deployed in high-stakes financial domains, yet they suffer from specific, reproducible hallucinations when performing arithmetic operations. Current mitigation strategies often treat the model…

Computation and Language · Computer Science 2025-12-01 Soham Mirajkar
‹ Prev 1 3 4 5 6 7 10 Next ›