9 papers · ranked by Valyu relevance
DeepReinforce Team, Xiaoya Li, Xiaofei Sun, Guoyin Wang + 3 more
Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms the best humans competitive programming: the most recent best result, Google's Gemini~3 Deep Think, attained 8th place even not being evaluated under live competition conditions. In…
Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain + 4 more
Competitive programming has become a rigorous benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). The International Olympiad in Informatics (IOI) stands out as one of the most prestigious annual competitions in competitive programming and has become a key benchmark…
Kaijian Zou, Xiong, Aaron, Yunxiang Zhang + 7 more
Competitive programming problems increasingly serve as valuable benchmarks to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as lack of exceptionally challenging problems, insufficient test case…
Tingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu + 1 more
Despite strong performance in competitive programming, the role of Large Language Models (LLMs) in supporting human learning in the same setting remains largely unexplored. In this work, we introduce UOJ-Bench, a benchmark designed to evaluate not only the problem-solving ability of LLMs, but also their ability to…
H. M. Shadman Tabib, Jaber Ahmed Deedar
Large Language Models (LLMs) have demonstrated impressive capabilities in natural language and code generation, and are increasingly deployed as automatic judges of model outputs and learning activities. Yet, their behavior on structured tasks such as predicting the difficulty of competitive programming problems…
Jörg Schneider, Sebastian Neef, Sebastian Koch
Security contests in the form of CTF (Capture The Flag) exercises are nowadays a common way to learn cyber security. 20 years ago at DIMVA 2006 the on-site CTF CIPHER II was one of the conference highlights and led to the foundation of the team ENOFLAG. In this poster, we reflect on the changes in the CTF gameplay and…
Maria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov + 4 more
LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks. By curating competitive programming problems, constantly adding fresh problems to the set, and filtering them by release dates, LCB provides contamination-aware evaluation and offers…
Authors not listed
Computational blind challenges offer critical, unbiased assessment opportunities to assess and accelerate scientific progress, as demonstrated by a breadth of breakthroughs over the last decade. We report the outcomes and key insights from an open science community blind challenge focused on computational methods in…
Ludwig Rappelt, Tim Wiedenmann, Steffen Held, Lars Heinke + 3 more
Introduction HYROX is a rapidly growing fitness format combining 8 km of running with 8 standardized workout stations. Despite its rapid growth, empirical evidence on performance correlates of success remains scarce. The purpose of this study is to investigate overall performance and its evolution as well as the…