OY Labs claims perfect ARC-AGI-3 public score with OY1-AGI
OY Labs, a Polish AI research company, has announced that its OY1-AGI system achieved a perfect 100.0% score on the ARC-AGI-3 public benchmark, solving all 183 levels across all 25 games. The result, published on 9 September 2026, used 6,732 actions and was completed in six hours and 48 minutes at a recorded API cost of $415.37. Two additional scorecards from 6 and 7 September document the same perfect score, with action counts of 6,659 and 6,714 respectively.
The company says OY1-AGI's leaderboard application remains under review by ARC Prize. OY Labs states that the three public scorecards are already available and that the result sits at the ceiling of the public scoring system, currently the highest score shown on the ARC-AGI-3 Community Leaderboard.
How the result was produced
The benchmark result was generated using OpenAI's GPT-6 Astra model at its High reasoning setting, guided by what OY Labs calls its OY1 harness, a proprietary scaffold that structures the model's problem-solving process. The company has open-sourced the benchmark implementation under the Apache 2.0 licence, including source code, reproduction instructions and results documentation.
It is worth being precise about what this result does and does not show. ARC Prize's own published figure for GPT-6 Astra with its Provider Adapter harness at High reasoning was 100% in 24 of 25 environments and 91.8% in one environment called TN36. OY Labs claims 100% across all 25, attributing the difference to its OY1 harness rather than the underlying model. The release notes explicitly that public games were used during development of the harness, meaning the result does not constitute a held-out evaluation. ARC Prize's verified semi-private score for the OpenAI configuration is 99.95% at a run cost of $18,817, a figure covering a different evaluation set and therefore not directly comparable to the $415.37 public-run figure.
Market context and competitive positioning
ARC-AGI benchmarks have become a focal point for AGI progress claims since ARC Prize, backed by François Chollet, introduced the challenge as a deliberate test of fluid intelligence rather than pattern memorisation. The competitive leaderboard has attracted entries from frontier-model labs, academic groups and an increasing number of well-funded startups. Retrodict, the next-ranked community entry below OY Labs, recorded 99.9% on the public set; NVIDIA's NOOA system achieved 85.1%.
OY Labs is a small, early-stage operation. The company says its next model, OY2-AGI, will embed a hard-coded ethical constitution into its design from the outset, with privacy protections and harm-prevention safeguards built in rather than added after training. Co-founder and chief executive Laurin Bylica described the company's goal as making "advanced intelligence affordable enough to use everywhere." Co-founder and CTO Loïc Bellez said the purpose of advancing human welfare should shape both capabilities and the limits built into them.
The company reports strong inbound interest from institutions seeking models trained to their specific requirements, though it has not named any customers, disclosed revenue or announced external funding. For now, OY Labs' commercial story rests almost entirely on its benchmark positioning and the cost efficiency suggested by its low API expenditure per run. Buyers and investors evaluating the result will reasonably want to see performance on held-out evaluations and independent replication before drawing strong conclusions about general-purpose AI capability.