English
Related papers

Related papers: A Note on How to Remove the $\ln\ln T$ Term from t…

200 papers

Post-training quantization (PTQ) has become a crucial tool for reducing the memory and compute costs of modern deep neural networks, including large language models (LLMs). Among PTQ algorithms, the OPTQ framework-also known as GPTQ-has…

Machine Learning · Computer Science 2026-04-13 Haoyu Zhang , Shihao Zhang , Ian Colbert , Rayan Saab

Large Language Models (LLMs) have demonstrated exceptional performance across a wide range of tasks, yet their significant computational and memory requirements present major challenges for deployment. A common approach uses Taylor…

Computation and Language · Computer Science 2026-03-10 Yijun Zhu , Jianxin Wang , Chengchao Shen

Recent theoretical work established the unsupervised identifiability of quantized factors under any diffeomorphism. The theory assumes that quantization thresholds correspond to axis-aligned discontinuities in the probability density of the…

Machine Learning · Computer Science 2025-11-27 Vitoria Barin-Pacela , Kartik Ahuja , Simon Lacoste-Julien , Pascal Vincent

In the hybrid kT-factorization formula, one initial-state parton momentum is space-like and carries non-vanishing transverse components, while the other is on-shell. We promote this factorization formula to next-to-leading order. Studying…

High Energy Physics - Phenomenology · Physics 2022-11-30 Andreas van Hameren , Leszek Motyka , Grzegorz Ziarko

This paper introduces KLong, an open-source LLM agent trained to solve extremely long-horizon tasks. The principle is to first cold-start the model via trajectory-splitting SFT, then scale it via progressive RL training. Specifically, we…

Artificial Intelligence · Computer Science 2026-04-28 Yue Liu , Yingwei Ma , Yibo Miao , Yanhao Li , Yuchong Xie , Xinlong Yang , Zhiyuan Hu , Flood Sung , Jiaheng Zhang , Bryan Hooi

In large-scale classification problems, the data set always be faced with frequent updates when a part of the data is added to or removed from the original data set. In this case, conventional incremental learning, which updates an existing…

Machine Learning · Computer Science 2021-01-15 Kaichen Zhou , Shiji Song , Gao Huang , Wu Cheng , Quan Zhou

We introduce a new criterion which tests if a given decomposition of a given ternary form $T$ of even degree is unique. The criterion is based on the analysis of the Hilbert function of the projective set of points $Z$ associated to the…

Algebraic Geometry · Mathematics 2020-07-21 Andrea Mazzon

Quantised neural networks (QNNs) shrink models and reduce inference energy through low-bit arithmetic, yet most still depend on a running statistics batch normalisation (BN) layer, preventing true integer-only deployment. Prior attempts…

Machine Learning · Computer Science 2025-12-19 Pengfei Sun , Wenyu Jiang , Piew Yoong Chee , Paul Devos , Dick Botteldooren

It has been a fascinating topic in the study of boundary layer theory about the well-posedness of Prandtl equation that was derived in 1904. Recently, new ideas about cancellation to overcome the loss of tangential derivatives were obtained…

Analysis of PDEs · Mathematics 2022-01-26 Tong Yang

We develop a modified online mirror descent framework that is suitable for building adaptive and parameter-free algorithms in unbounded domains. We leverage this technique to develop the first unconstrained online linear optimization…

Machine Learning · Computer Science 2024-02-12 Andrew Jacobsen , Ashok Cutkosky

This work develops a zero-shot mechanism, Comp-LTL, for an agent to satisfy a Linear Temporal Logic (LTL) specification given existing task primitives trained via reinforcement learning (RL). Autonomous robots often need to satisfy spatial…

Robotics · Computer Science 2024-12-17 Taylor Bergeron , Zachary Serlin , Kevin Leahy

We consider the problem of prediction with expert advice when the losses of the experts have low-dimensional structure: they are restricted to an unknown $d$-dimensional subspace. We devise algorithms with regret bounds that are independent…

Machine Learning · Computer Science 2016-05-24 Elad Hazan , Tomer Koren , Roi Livni , Yishay Mansour

We propose a linear contextual bandit algorithm with $O(\sqrt{dT\log T})$ regret bound, where $d$ is the dimension of contexts and $T$ isthe time horizon. Our proposed algorithm is equipped with a novel estimator in which exploration is…

Machine Learning · Statistics 2023-03-30 Wonyoung Kim , Myunghee Cho Paik , Min-hwan Oh

We derive a tight generalization bound for quantum machine learning that is applicable to a wide range of supervised tasks, data, and models. Our bound is both efficiently computable and free of big-O notation. Furthermore, we point out…

Quantum Physics · Physics 2025-10-29 Xin Wang , Rebing Wu

The scaling of model and data sizes has reshaped the AI landscape, establishing finetuning pretrained models as the standard paradigm for solving downstream tasks. However, dominant finetuning methods typically rely on weight adaptation,…

Machine Learning · Computer Science 2026-01-16 Leyang Hu , Matteo Gamba , Randall Balestriero

We give a simple optimistic algorithm for which it is easy to derive regret bounds of $\tilde{O}(\sqrt{t_{\rm mix} SAT})$ after $T$ steps in uniformly ergodic Markov decision processes with $S$ states, $A$ actions, and mixing time parameter…

Machine Learning · Computer Science 2019-01-23 Ronald Ortner

Combining model-based and model-free reinforcement learning approaches, this paper proposes and analyzes an $\epsilon$-policy gradient algorithm for the online pricing learning task. The algorithm extends $\epsilon$-greedy algorithm by…

Machine Learning · Computer Science 2024-05-07 Lukasz Szpruch , Tanut Treetanthiploet , Yufei Zhang

By extending the Johnson--Barron projection method from one dimension to high dimensions and utilizing a Wang type dimension-free Harnack inequality, we obtain a new quantitative bound for the entropic central limit theorem under the…

Probability · Mathematics 2026-04-08 Chang-song Deng , Lin Wang , Lihu Xu

We produce a series of Central Limit Theorems (CLTs) associated to compact metric measure spaces $(K,d,\eta)$, with $\eta$ a reasonable probability measure. For the first CLT, we can ignore $\eta$ by isometrically embedding $K$ into…

Probability · Mathematics 2020-01-14 Steven Rosenberg , Jie Xu

We improve the Solovay--Kitaev theorem and algorithm for a general finite, inverse-closed generating set acting on a qudit. Prior versions of the algorithm efficiently find a word of length $O(n^{3+\delta})$ to approximate an arbitrary…

Quantum Physics · Physics 2025-10-09 Greg Kuperberg