English
Related papers

Related papers: Perfect score on IPhO 2025 theory by Gemini agent

200 papers

Physics provides fundamental laws that describe and predict the natural world. AI systems aspiring toward more general, real-world intelligence must therefore demonstrate strong physics problem-solving abilities: to formulate and apply…

Artificial Intelligence · Computer Science 2025-09-03 Jiahao Qiu , Jingzhe Shi , Xinzhe Juan , Zelin Zhao , Jiayi Geng , Shilong Liu , Hongru Wang , Sanfeng Wu , Mengdi Wang

The International Mathematical Olympiad (IMO) is widely regarded as the world championship of high-school mathematics. IMO problems are renowned for their difficulty and novelty, demanding deep insight, creativity, and rigor. Although large…

Artificial Intelligence · Computer Science 2025-10-01 Yichen Huang , Lin F. Yang

The International Mathematical Olympiad (IMO) is perhaps the most celebrated mental competition in the world and as such is among the greatest grand challenges for Artificial Intelligence (AI). The IMO Grand Challenge, recently formulated,…

Logic in Computer Science · Computer Science 2020-11-02 Filip Marić , Sana Stojanović-{\Dj}urđević

Physics is central to understanding and shaping the real world, and the ability to solve physics problems is a key indicator of real-world physical intelligence. Physics Olympiads, renowned as the crown of competitive physics, provide a…

Artificial Intelligence · Computer Science 2025-09-30 Fangchen Yu , Junchi Yao , Ziyi Wang , Haiyuan Wan , Youling Huang , Bo Zhang , Shuyue Hu , Dongzhan Zhou , Ning Ding , Ganqu Cui , Lei Bai , Wanli Ouyang , Peng Ye

Olympiad-level physics problem-solving significantly challenges both humans and artificial intelligence (AI), as it requires integrating appropriate modeling, application of physical principles, and precise calculation within long reasoning…

Recently, the physical capabilities of (M)LLMs have garnered increasing attention. However, existing benchmarks for physics suffer from two major gaps: they neither provide systematic and up-to-date coverage of real-world physics…

Recent progress in large language models (LLMs) has moved the frontier from puzzle-solving to science-grade reasoning-the kind needed to tackle problems whose answers must stand against nature, not merely fit a rubric. Physics is the…

OpenAI and DeepMind's AIs recently got gold at the IMO math olympiad and ICPC programming competition. We show frontier AI is similarly good at hacking by letting GPT-5 compete in elite CTF cybersecurity competitions. In one of this year's…

Cryptography and Security · Computer Science 2025-11-17 Reworr , Artem Petrov , Dmitrii Volkov

Opacity is a general framework modeling security properties of systems interacting with a passive attacker. Initial-and-final-state opacity (IFO) generalizes the classical notions of opacity, such as current-state opacity and initial-state…

Formal Languages and Automata Theory · Computer Science 2024-12-25 Tomáš Masopust , Petr Osička

Recent math benchmarks for large language models (LLMs) such as MathArena indicate that state-of-the-art reasoning models achieve impressive performance on mathematical competitions like AIME, with the leading model, Gemini-2.5-Pro,…

As large language models (LLMs) reach high scores on established mathematical benchmarks, such as GSM8K and MATH, the research community has turned to International Mathematical Olympiad (IMO) problems to push the evaluation frontier.…

Artificial Intelligence · Computer Science 2025-09-10 Ziye Chen , Chengwei Qin , Yao Shu

In this report, we pose the following question: Who is the most intelligent AI model to date, as measured by the OlympicArena (an Olympic-level, multi-discipline, multi-modal benchmark for superintelligent AI)? We specifically focus on the…

Computation and Language · Computer Science 2024-06-27 Zhen Huang , Zengzhi Wang , Shijie Xia , Pengfei Liu

Olympiad-level benchmarks in mathematics and physics are crucial testbeds for advanced AI reasoning, but chemistry, with its unique multimodal symbolic language, has remained an open challenge. We introduce ChemO, a new benchmark built from…

Artificial Intelligence · Computer Science 2025-12-10 Qiang Xu , Shengyuan Bai , Leqing Chen , Zijing Liu , Yu Li

Artificial intelligence (AI) is poised to transform education, but the research community lacks a robust, general benchmark to evaluate AI models for learning. To assess state-of-the-art support for educational use cases, we ran an "arena…

Between 2021 and 2023, AI-Olympics, a series of online AI competitions was hosted by the online evaluation platform Jidi in collaboration with the IJCAI committee. In these competitions, an agent is required to accomplish diverse sports…

Multiagent Systems · Computer Science 2024-05-24 Chen Wang , Yan Song , Shuai Wu , Sa Wu , Ruizhi Zhang , Shu Lin , Haifeng Zhang

Proprietary AI systems have recently demonstrated impressive capabilities on complex proof-based problems, with gold-level performance reported at the 2025 International Mathematical Olympiad (IMO). However, the training pipelines behind…

Artificial Intelligence · Computer Science 2026-04-07 LM-Provers , Yuxiao Qu , Amrith Setlur , Jasper Dekoninck , Edward Beeching , Jia Li , Ian Wu , Lewis Tunstall , Aviral Kumar

State-of-the-art (SOTA) LLMs have progressed from struggling on proof-based Olympiad problems to solving most of the IMO 2025 problems, with leading systems reportedly handling 5 of 6 problems. Given this progress, we assess how well these…

We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate…

The Young Physicists Tournament is an established team-oriented scientific competition between high school students from 37 countries on 5 continents. The competition consists of scientific discussions called Fights. Three or four teams…

Artificial Intelligence · Computer Science 2021-04-20 Katarína Cechlárová , Ágnes Cseh , Zsuzsanna Jankó , Marián Kireš , Lukáš Miňo

Finding the right north-star metrics is highly critical for advancing the mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focus on getting correct short…

‹ Prev 1 2 3 10 Next ›