English
Related papers

Related papers: Logical Randomized Benchmarking

200 papers

Reasoning has emerged as the next major frontier for language models (LMs), with rapid advances from both academic and industrial labs. However, this progress often outpaces methodological rigor, with many evaluations relying on…

Machine Learning · Computer Science 2025-10-08 Andreas Hochlehnert , Hardik Bhatnagar , Vishaal Udandarao , Samuel Albanie , Ameya Prabhu , Matthias Bethge

We introduce the lookahead-bounded Q-learning (LBQL) algorithm, a new, provably convergent variant of Q-learning that seeks to improve the performance of standard Q-learning in stochastic environments through the use of ``lookahead'' upper…

Machine Learning · Computer Science 2020-06-30 Ibrahim El Shar , Daniel R. Jiang

As we approach the era of quantum advantage, when quantum computers (QCs) can outperform any classical computer on particular tasks, there remains the difficult challenge of how to validate their performance. While algorithmic success can…

We apply quantum error mitigation techniques to a variety of benchmark problems and quantum computers to evaluate the performance of quantum error mitigation in practice. To do so, we define an empirically motivated, resource-normalized…

Quantum Physics · Physics 2023-08-21 Vincent Russo , Andrea Mari , Nathan Shammah , Ryan LaRose , William J. Zeng

Formal mathematical reasoning remains a critical challenge for artificial intelligence, hindered by limitations of existing benchmarks in scope and scale. To address this, we present FormalMATH, a large-scale Lean4 benchmark comprising…

Quantum error correction provides a path to reach practical quantum computing by combining multiple physical qubits into a logical qubit, where the logical error rate is suppressed exponentially as more qubits are added. However, this…

Quantum Physics · Physics 2025-04-08 Rajeev Acharya , Laleh Aghababaie-Beni , Igor Aleiner , Trond I. Andersen , Markus Ansmann , Frank Arute , Kunal Arya , Abraham Asfaw , Nikita Astrakhantsev , Juan Atalaya , Ryan Babbush , Dave Bacon , Brian Ballard , Joseph C. Bardin , Johannes Bausch , Andreas Bengtsson , Alexander Bilmes , Sam Blackwell , Sergio Boixo , Gina Bortoli , Alexandre Bourassa , Jenna Bovaird , Leon Brill , Michael Broughton , David A. Browne , Brett Buchea , Bob B. Buckley , David A. Buell , Tim Burger , Brian Burkett , Nicholas Bushnell , Anthony Cabrera , Juan Campero , Hung-Shen Chang , Yu Chen , Zijun Chen , Ben Chiaro , Desmond Chik , Charina Chou , Jahan Claes , Agnetta Y. Cleland , Josh Cogan , Roberto Collins , Paul Conner , William Courtney , Alexander L. Crook , Ben Curtin , Sayan Das , Alex Davies , Laura De Lorenzo , Dripto M. Debroy , Sean Demura , Michel Devoret , Agustin Di Paolo , Paul Donohoe , Ilya Drozdov , Andrew Dunsworth , Clint Earle , Thomas Edlich , Alec Eickbusch , Aviv Moshe Elbag , Mahmoud Elzouka , Catherine Erickson , Lara Faoro , Edward Farhi , Vinicius S. Ferreira , Leslie Flores Burgos , Ebrahim Forati , Austin G. Fowler , Brooks Foxen , Suhas Ganjam , Gonzalo Garcia , Robert Gasca , Élie Genois , William Giang , Craig Gidney , Dar Gilboa , Raja Gosula , Alejandro Grajales Dau , Dietrich Graumann , Alex Greene , Jonathan A. Gross , Steve Habegger , John Hall , Michael C. Hamilton , Monica Hansen , Matthew P. Harrigan , Sean D. Harrington , Francisco J. H. Heras , Stephen Heslin , Paula Heu , Oscar Higgott , Gordon Hill , Jeremy Hilton , George Holland , Sabrina Hong , Hsin-Yuan Huang , Ashley Huff , William J. Huggins , Lev B. Ioffe , Sergei V. Isakov , Justin Iveland , Evan Jeffrey , Zhang Jiang , Cody Jones , Stephen Jordan , Chaitali Joshi , Pavol Juhas , Dvir Kafri , Hui Kang , Amir H. Karamlou , Kostyantyn Kechedzhi , Julian Kelly , Trupti Khaire , Tanuj Khattar , Mostafa Khezri , Seon Kim , Paul V. Klimov , Andrey R. Klots , Bryce Kobrin , Pushmeet Kohli , Alexander N. Korotkov , Fedor Kostritsa , Robin Kothari , Borislav Kozlovskii , John Mark Kreikebaum , Vladislav D. Kurilovich , Nathan Lacroix , David Landhuis , Tiano Lange-Dei , Brandon W. Langley , Pavel Laptev , Kim-Ming Lau , Loïck Le Guevel , Justin Ledford , Kenny Lee , Yuri D. Lensky , Shannon Leon , Brian J. Lester , Wing Yan Li , Yin Li , Alexander T. Lill , Wayne Liu , William P. Livingston , Aditya Locharla , Erik Lucero , Daniel Lundahl , Aaron Lunt , Sid Madhuk , Fionn D. Malone , Ashley Maloney , Salvatore Mandrá , Leigh S. Martin , Steven Martin , Orion Martin , Cameron Maxfield , Jarrod R. McClean , Matt McEwen , Seneca Meeks , Anthony Megrant , Xiao Mi , Kevin C. Miao , Amanda Mieszala , Reza Molavi , Sebastian Molina , Shirin Montazeri , Alexis Morvan , Ramis Movassagh , Wojciech Mruczkiewicz , Ofer Naaman , Matthew Neeley , Charles Neill , Ani Nersisyan , Hartmut Neven , Michael Newman , Jiun How Ng , Anthony Nguyen , Murray Nguyen , Chia-Hung Ni , Thomas E. O'Brien , William D. Oliver , Alex Opremcak , Kristoffer Ottosson , Andre Petukhov , Alex Pizzuto , John Platt , Rebecca Potter , Orion Pritchard , Leonid P. Pryadko , Chris Quintana , Ganesh Ramachandran , Matthew J. Reagor , David M. Rhodes , Gabrielle Roberts , Eliott Rosenberg , Emma Rosenfeld , Pedram Roushan , Nicholas C. Rubin , Negar Saei , Daniel Sank , Kannan Sankaragomathi , Kevin J. Satzinger , Henry F. Schurkus , Christopher Schuster , Andrew W. Senior , Michael J. Shearn , Aaron Shorter , Noah Shutty , Vladimir Shvarts , Shraddha Singh , Volodymyr Sivak , Jindra Skruzny , Spencer Small , Vadim Smelyanskiy , W. Clarke Smith , Rolando D. Somma , Sofia Springer , George Sterling , Doug Strain , Jordan Suchard , Aaron Szasz , Alex Sztein , Douglas Thor , Alfredo Torres , M. Mert Torunbalci , Abeer Vaishnav , Justin Vargas , Sergey Vdovichev , Guifre Vidal , Benjamin Villalonga , Catherine Vollgraff Heidweiller , Steven Waltman , Shannon X. Wang , Brayden Ware , Kate Weber , Theodore White , Kristi Wong , Bryan W. K. Woo , Cheng Xing , Z. Jamie Yao , Ping Yeh , Bicheng Ying , Juhwan Yoo , Noureldin Yosri , Grayson Young , Adam Zalcman , Yaxing Zhang , Ningfeng Zhu , Nicholas Zobrist

Benchmarking the performance of quantum optimization algorithms is crucial for identifying utility for industry-relevant use cases. Benchmarking processes vary between optimization applications and depend on user-specified goals. The…

Complex reasoning tasks often rely on the ability to consistently and accurately apply simple rules across incremental steps, a foundational capability which we term "level-0" reasoning. To systematically evaluate this capability, we…

Programming Languages · Computer Science 2025-04-14 Simeng Sun , Cheng-Ping Hsieh , Faisal Ladhak , Erik Arakelyan , Santiago Akle Serano , Boris Ginsburg

The overall goal of this paper is to investigate the theoretical foundations of algorithmic verification techniques for first order linear logic specifications. The fragment of linear logic we consider in this paper is based on the linear…

Programming Languages · Computer Science 2007-05-23 M. Bozzano , G. Delzanno , M. Martelli

Probabilistic quantum error correction is an error-correcting procedure which uses postselection to determine if the encoded information was successfully restored. In this work, we deeply analyze probabilistic version of the…

Quantum Physics · Physics 2023-04-12 Ryszard Kukulski , Łukasz Pawela , Zbigniew Puchała

The realization of quantum error correction is an essential ingredient for reaching the full potential of fault-tolerant universal quantum computation. Using a range of different schemes, logical qubits can be redundantly encoded in a set…

We present an experimental procedure to determine the usefulness of a measurement scheme for quantum error correction (QEC). A QEC scheme typically requires the ability to prepare entangled states, to carry out multi-qubit measurements, and…

Quantum Physics · Physics 2015-06-04 Gabrielle Denhez , Alexandre Blais , David Poulin

Large Language Models (LLMs) are increasingly being used to automate programming tasks. Yet, LLMs' capabilities in reasoning about program semantics are still inadequately studied, leaving significant potential for further exploration. This…

Programming Languages · Computer Science 2025-05-30 Thanh Le-Cong , Bach Le , Toby Murray

The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluation of toxicity benchmarks. As organizations increasingly rely on these benchmarks to…

Artificial Intelligence · Computer Science 2026-05-12 Regina Gugg , Selina Niederländer , Andreas Stöckl , Martin Flechl

Logical reasoning is a fundamental aspect of human intelligence and an essential capability for multimodal large language models (MLLMs). Despite the significant advancement in multimodal reasoning, existing benchmarks fail to…

Artificial Intelligence · Computer Science 2025-05-28 Jiakang Yuan , Tianshuo Peng , Yilei Jiang , Yiting Lu , Renrui Zhang , Kaituo Feng , Chaoyou Fu , Tao Chen , Lei Bai , Bo Zhang , Xiangyu Yue

The critique capacity of Large Language Models (LLMs) is essential for reasoning abilities, which can provide necessary suggestions (e.g., detailed analysis and constructive feedback). Therefore, how to evaluate the critique capacity of…

In previous work, we proposed a method for leveraging efficient classical simulation algorithms to aid in the analysis of large-scale fault tolerant circuits implemented on hypothetical quantum information processors. Here, we extend those…

Quantum Physics · Physics 2014-02-12 Daniel Puzzuoli , Christopher Granade , Holger Haas , Ben Criger , Easwar Magesan , D. G. Cory

Utility-scale quantum computers require quantum error correcting codes with large numbers of physical qubits to achieve sufficiently low logical error rates. The performance of quantum error correction (QEC) is generally predicted through…

Integrating logic rules with other language features is increasingly sought after for advanced applications that require knowledge-base capabilities. To address this demand, increasingly more languages and extensions for such integration…

Programming Languages · Computer Science 2023-08-31 Yanhong A. Liu , Scott D. Stoller , Yi Tong , K. Tuncay Tekle

The emphasis is made on the juxtaposition of (quantum~theorem) proving versus quantum (theorem~proving). The logical contents of verification of the statements concerning quantum systems is outlined. The Zittereingang (trembling input)…

Quantum Physics · Physics 2009-10-28 R. R. Zapatrin