The result
In the ResearchClawBench leaderboard snapshot dated August 19, 2026, OpenEvo is listed in first place with an overall score of 32.7
The run is recorded as a community submission using GPT-5.5, with a reported cost of $2.3 and a runtime of 10 minutes 19 seconds
Why this matters to us
This result is a useful checkpoint for EvoLab’s central idea: agent improvement should remain connected to real research tasks, observable execution, and reusable experience
A leaderboard is one measurement rather than a final verdict, but the result gives us a concrete baseline for evaluating what the system learns and where it still needs to improve
What comes next
We will continue testing OpenEvo across scientific domains, studying failure cases, and turning validated task experience into better memory, skills, and agent systems for later sessions