Related papers: Conformal Bootstrap with Reinforcement Learning
Reinforcement finetuning (RFT) is a key technique for aligning Large Language Models (LLMs) with human preferences and enhancing reasoning, yet its effectiveness is highly sensitive to which tasks are explored during training. Uniform task…
Reinforcement learning (RL) is often credited with improving language model reasoning and generalization at the expense of degrading memorized knowledge. We challenge this narrative by observing that RL-enhanced models consistently…
The computational cost of stiff chemical kinetics remains a dominant bottleneck in reacting-flow simulation, yet hybrid integration strategies are typically driven by hand-tuned heuristics or supervised predictors that make myopic decisions…
Finite-size effects limit the accuracy with which conformal data can be extracted from lattice simulations of critical systems. While action improvement suppresses some corrections to scaling, it does not address operator-dependent effects…
Mobile robots are increasingly being employed for performing complex tasks in dynamic environments. Reinforcement learning (RL) methods are recognized to be promising for specifying such tasks in a relatively simple manner. However, the…
Reinforcement learning (RL) is a fundamental framework for sequential decision-making, in which an agent learns an optimal policy through interactions with an unknown environment. In settings with function approximation, many existing RL…
We demonstrate that the Ising model on a general triangular graph with 3 distinct couplings $K_1,K_2,K_3$ corresponds to an affine transformed conformal field theory (CFT). Full conformal invariance of the $c= 1/2$ minimal CFT is restored…
Quadratic programming is a workhorse of modern nonlinear optimization, control, and data science. Although regularized methods offer convergence guarantees under minimal assumptions on the problem data, they can exhibit the slow…
We use the embedding formalism to construct conformal fields in $D$ dimensions, by restricting Lorentz-invariant ensembles of homogeneous neural networks in $(D+2)$ dimensions to the projective null cone. Conformal correlators may be…
In this thesis we study two-dimensional conformal field theories with Virasoro algebra symmetry, following the conformal bootstrap approach. Under the assumption that degenerate fields exist, we provide an extension of the analytic…
Transformers can acquire Chain-of-Thought (CoT) capabilities to solve complex reasoning tasks through fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary approaches to this end. In this work, we…
We present a systematic exploration of conformal field theories (CFTs) constrained by duality-inspired fusion rules using the conformal bootstrap. We classify the operator spectrum into three sectors: $[\sigma]$, $[\epsilon]$, and $[1]$.…
Adapting large language models to multiple tasks can cause cross-skill interference, where improvements for one skill degrade another. While methods such as LoRA impose orthogonality constraints at the weight level, they do not fully…
We use the conformal bootstrap to perform a precision study of 3d maximally supersymmetric ($\mathcal{N}=8$) SCFTs that describe the IR physics on $N$ coincident M2-branes placed either in flat space or at a $\C^4/\Z_2$ singularity. First,…
Geosteering, a key component of drilling operations, traditionally involves manual interpretation of various data sources such as well-log data. This introduces subjective biases and inconsistent procedures. Academic attempts to solve…
We explain how the axioms of Conformal Field Theory are used to make predictions about critical exponents of continuous phase transitions in three dimensions, via a procedure called the conformal bootstrap. The method assumes conformal…
Reinforcement learning (RL) has proven to be well-performed and general-purpose in the inventory control (IC). However, further improvement of RL algorithms in the IC domain is impeded due to two limitations of online experience. First,…
This paper studies satisfaction of temporal properties on unknown stochastic processes that have continuous state spaces. We show how reinforcement learning (RL) can be applied for computing policies that are finite-memory and deterministic…
The dimensional reductions in the branched polymer and the random field Ising model (RFIM) are discussed by a conformal bootstrap method. The small size minors are applied for the evaluations of the scale dimensions of these two models and…
Online matching problems arise in many complex systems, from cloud services and online marketplaces to organ exchange networks, where timely, principled decisions are critical for maintaining high system performance. Traditional heuristics…