Related papers: Term Orders for Optimistic Lambda-Superposition
Learning to Optimize (L2O) approaches, including algorithm unrolling, plug-and-play methods, and hyperparameter learning, have garnered significant attention and have been successfully applied to the Alternating Direction Method of…
Neural combinatorial optimization (NCO) is a promising learning-based approach for solving challenging combinatorial optimization problems without specialized algorithm design by experts. However, most constructive NCO methods cannot solve…
Higher-order QED radiative corrections to muon decay spectrum are evaluated within the QED structure function approach in the next-to-leading order logarithmic approximation. New analytical results are given in the…
In earlier work with C.~Monical, we introduced the notion of a K-crystal, with applications to K-theoretic Schubert calculus and the study of Lascoux polynomials. We conjectured that such a K-crystal structure existed on the set of…
Comparison-Based Optimization (CBO) is an optimization paradigm that assumes only very limited access to the objective function f(x). Despite the growing relevance of CBO to real-world applications, this field has received little attention…
We compare three finite element based methods designed for two-sided bounds of eigenvalues of symmetric elliptic second order operators. The first method is known as the Lehmann-Goerisch method. The second method is based on…
Many applications of large language models (LLMs), ranging from chatbots to creative writing, require nuanced subjective judgments that can differ significantly across different groups. Existing alignment algorithms can be expensive to…
In this paper we study a class of constrained minimax problems. In particular, we propose a first-order augmented Lagrangian method for solving them, whose subproblems turn out to be a much simpler structured minimax problem and are…
We present a novel quantum optimization-based route compression technique that significantly reduces storage requirements compared to conventional methods. Route optimization systems face critical challenges in efficiently storing selected…
Direct Preference Optimization is an offline post-SFT method for aligning language models from preference pairs, with strong results in instruction following and summarization. However, DPO's sequence-level implicit reward can be brittle…
This paper begins by extending the notion of a combinatorial configuration of points and lines to a combinatorial configuration of points and planes that we refer to as configurations of order $2$. We then proceed to investigate a further…
We find two series expansions for Legendre's second incomplete elliptic integral $E(\lambda, k)$ in terms of recursively computed elementary functions. Both expansions converge at every point of the unit square in the $(\lambda, k)$ plane.…
Large Language Models (LLMs) have demonstrated remarkable potential in automating software development tasks. While recent advances leverage Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) to align models with human…
We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in aligning language models, eliminating the need for reward…
We present a new methodology for decomposing flows with multiple transports that further extends the shifted proper orthogonal decomposition (sPOD). The sPOD tries to approximate transport-dominated flows by a sum of co-moving data fields.…
In real-world services such as ChatGPT, aligning models based on user feedback is crucial for improving model performance. However, due to the simplicity and convenience of providing feedback, users typically offer only basic binary…
Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce Continuous Utility Direct Preference…
Large Language Models (LLMs) have become increasingly popular due to their ability to process and generate natural language. However, as they are trained on massive datasets of text, LLMs can inherit harmful biases and produce outputs that…
We show the NP-completeness of the existential theory of term algebras with the Knuth-Bendix order by giving a nondeterministic polynomial-time algorithm for solving Knuth-Bendix ordering constraints.
In this paper, we introduce a Kantorovich version of the Bernstein-type logarithmic operators. The idea comes from the wide literature concerning exponential polynomials that preserve exponential functions: here, the exponential weights are…