English
Related papers

Related papers: Perfect score on IPhO 2025 theory by Gemini agent

200 papers

In this paper we summarize the results of the Putnam-like benchmark published by Google DeepMind. This dataset consists of 96 original problems in the spirit of the Putnam Competition and 576 solutions generated by LLMs. We analyze the…

Machine Learning · Computer Science 2026-02-03 Bartosz Bieganowski , Daniel Strzelecki , Robert Skiba , Mateusz Topolewski

The theory of algorithmic fair allocation is within the center of multi-agent systems and economics in the last decade due to its industrial and social importance. At a high level, the problem is to assign a set of items that are either…

Computer Science and Game Theory · Computer Science 2022-02-18 Haris Aziz , Bo Li , Herve Moulin , Xiaowei Wu

It is often advantageous to train models on a subset of the available train examples, because the examples are of variable quality or because one would like to train with fewer examples, without sacrificing performance. We present Gradient…

Machine Learning · Computer Science 2024-07-30 Dante Everaert , Christopher Potts

The first Global e-Competition on Astronomy and Astrophysics was held online in September-October 2020 as a replacement for the International Olympiad on Astronomy and Astrophysics, which was postponed due to the COVID-19 pandemic. Despite…

Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 Deep Think,…

Artificial Intelligence · Computer Science 2026-04-06 DeepReinforce Team , Xiaoya Li , Xiaofei Sun , Guoyin Wang , Songqiao Su , Chris Shum , Jiwei Li

The Keynesian Beauty Contest is a classical game in which strategic agents seek to both accurately guess the true state of the world as well as the average action of all agents. We study an augmentation of this game where agents are…

Computer Science and Game Theory · Computer Science 2019-05-03 Hadi Elzayn , Zachary Schutzman

Generative AI has the potential to transform personalization and accessibility of education. However, it raises serious concerns about accuracy and helping students become independent critical thinkers. In this study, we designed a helpful…

How can we enable machines to make sense of the world, and become better at learning? To approach this goal, I believe viewing intelligence in terms of many integral aspects, and also a universal two-term tradeoff between task performance…

Machine Learning · Computer Science 2020-01-22 Tailin Wu

Autonomous systems such as self-driving cars and general-purpose robots are safety-critical systems that operate in highly uncertain and dynamic environments. We propose an interactive multi-agent framework where the system-under-design is…

Machine Learning · Computer Science 2021-07-07 Xin Qin , Nikos Aréchiga , Andrew Best , Jyotirmoy Deshmukh

Static benchmarks capture only part of how large language models behave in practice. Real systems place models inside repeated loops with time limits, formatting constraints, and failure modes. We study this setting in a timed multi-phase…

Artificial Intelligence · Computer Science 2026-05-22 H. C. Ekne

Most recently developed approaches to cooperative multi-agent reinforcement learning in the \emph{centralized training with decentralized execution} setting involve estimating a centralized, joint value function. In this paper, we…

As a fundraising method, Initial Coin Offering (ICO) has raised billions of dollars for thousands of startups in the past two years. Existing ICO mechanisms place more emphasis on the short-term benefits of maximal fundraising while…

Computer Science and Game Theory · Computer Science 2020-02-27 Mingyu Guo , Zhenghui Wang , Yuko Sakurai

This report summarizes IROS 2019-Lifelong Robotic Vision Competition (Lifelong Object Recognition Challenge) with methods and results from the top $8$ finalists (out of over~$150$ teams). The competition dataset (L)ifel(O)ng (R)obotic…

In 2021 the Johns Hopkins University Applied Physics Laboratory held an internal challenge to develop artificially intelligent (AI) agents that could excel at the collaborative card game Hanabi. Agents were evaluated on their ability to…

Artificial Intelligence · Computer Science 2021-11-19 Nicholas Kantack

The first decade of this century has seen the nascency of the first mathematical theory of general artificial intelligence. This theory of Universal Artificial Intelligence (UAI) has made significant contributions to many theoretical,…

Artificial Intelligence · Computer Science 2013-05-17 Marcus Hutter

The rapid evolution of GUI-enabled agents has rendered traditional CAPTCHAs obsolete. While previous benchmarks like OpenCaptchaWorld established a baseline for evaluating multimodal agents, recent advancements in reasoning-heavy models,…

Machine Learning · Computer Science 2026-02-10 Jiacheng Liu , Yaxin Luo , Jiacheng Cui , Xinyi Shang , Xiaohan Zhao , Zhiqiang Shen

We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focuses on four core principles related to macroscopic objects:…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Florian Bordes , Quentin Garrido , Justine T Kao , Adina Williams , Michael Rabbat , Emmanuel Dupoux

Humans interact with objects all the time. Enabling a humanoid to learn human-object interaction (HOI) is a key step for future smart animation and intelligent robotics systems. However, recent progress in physics-based HOI requires…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Yinhuai Wang , Jing Lin , Ailing Zeng , Zhengyi Luo , Jian Zhang , Lei Zhang

This study compared the classification performance of Gemini Pro and GPT-4V in educational settings. Employing visual question answering (VQA) techniques, the study examined both models' abilities to read text-based rubrics and then…

Artificial Intelligence · Computer Science 2024-01-18 Gyeong-Geon Lee , Ehsan Latif , Lehong Shi , Xiaoming Zhai

Proximal policy optimization (PPO) algorithm is a deep reinforcement learning algorithm with outstanding performance, especially in continuous control tasks. But the performance of this method is still affected by its exploration ability.…

Machine Learning · Computer Science 2020-11-12 Junwei Zhang , Zhenghao Zhang , Shuai Han , Shuai Lü
‹ Prev 1 3 4 5 6 7 10 Next ›