Related papers: Corrigendum to "Managerial Incentive Problems: A D…
We correct a mistake in the paper ["On weighted iterated Hardy-type inequalities", Positivity, 22 (1) (2018), 275-299]. -- In this paper the inequality $$ \bigg( \int_0^{\infty} \bigg( \int_x^{\infty} \bigg( \int_t^{\infty} h \bigg)^q…
While recent advances have boosted LM proficiency in linguistic benchmarks, LMs consistently struggle to reason correctly on complex tasks like mathematics. We turn to Reinforcement Learning from Human Feedback (RLHF) as a method with which…
We consider a convex optimization problem with many linear inequality constraints. To deal with a large number of constraints, we provide a penalty reformulation of the problem, where the penalty is a variant of the one-sided Huber loss…
The tendency of repeating past choices more often than expected from the history of outcomes has been repeatedly empirically observed in reinforcement learning experiments. It can be explained by at least two computational processes:…
A gap in the proof of Theorem 3.5 in the paper ``A new iteration process for approximation of common fixed points for finite families of total asymtotically nonexpansive mappings". Int. J. Math. Math. Sci. vol. 2009,…
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of…
We survey and unify results on elimination of dominated strategies by monotonic dynamics and prove some new results that may be seen as dual to those of Hofbauer and Weibull (J. Econ. Theory, 1996, 558-573) on convex monotonic dynamics.
A correction to the specification of the mechanism proposed in "An Efficient Game Form for Unicast Service Provisioning" is given.
We correct a mistake in the paper "On the cuspidal cohomology of $S$-arithmetic subgroups of reductive groups over number fields " by A. Borel, J.-P. Labesse, J. Schwermer that appeared in Compositio Mathematica, 102 (1996), 1-40.
This is a resubmission of preprint 9401008 , which has some TeXnical errors introduced by the "reform" procedure (designed to avoid precisely these problems!). The original can be formatted by editing out the messages "%% following line…
Rejoinder of "Instrumental Variables: An Econometrician's Perspective" by Guido W. Imbens [arXiv:1410.0163].
This paper contains material originally presented in the first submission of the paper "On the Optimal Control of Impulsive Hybrid Systems On Riemannian Manifolds" to SIAM Journal on Control and Optimization on the 28th of February, 2012.…
Comment on ``Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data'' [arXiv:0804.2958]
Comment on ``Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data'' [arXiv:0804.2958]
The multistage stochastic variational inequality is reformulated into a variational inequality with separable structure through introducing a new variable. The prediction-correction ADMM which was originally proposed in [B.-S. He, L.-Z.…
This is a comment on Economic Letters DOI http://dx.doi.org/10.1016/j.econlet.2015.10.015. We show that due to some methodological aspects the main conclusions of the above mentioned paper should be a little bit altered.
Reinforcement Learning with Verifiable Rewards (RLVR) improves final-answer accuracy on reasoning tasks, but it does not reliably improve reasoning quality. Because outcome rewards only assess final answers, they also reward spurious…
We provide novel theoretical results regarding local optima of regularized $M$-estimators, allowing for nonconvexity in both loss and penalty functions. Under restricted strong convexity on the loss and suitable regularity conditions on the…
This manuscript, a revised version of arXiv:0811.3168v1, was inadvertently submitted as a separate paper. It can now be accessed, including some final corrections for the published version, as arXiv:0811.3168v2.
In [Ritika Garg et al., Phys. Rev. C 100, 069901(E) (2019)] the experimental results on the polarization asysmetry were revised due to a claimed change of the geometry asymmetry. However, the revised results can not be reproduced as claimed…