Reading the AutomationBench results
The reported results, sample scope and limits on cross-model comparisons.
This article retains the frozen evaluation previously shown on the homepage. It describes one evaluation, not a promise about future task success, cost or speed.
The current 100-case sample and the official held-out references use different sets. Reference values are not a cross-model ranking. Read the sample scope, strict-pass criterion and original evaluation evidence alongside the result.
Measured result
Mission raises the same Luna to 34.00%
100 AutomationBench cases scored by strict pass criteria: the unassisted model first, then the result after full OpenCorvus execution.
AutomationBench · Strict pass rate
OpenCorvus Mission Base
- Evaluated cases
- 100
- Current frozen sample
- Absolute lift
- +25.93 pp
- percentage points
- Versus original Luna
- 4.21×
- strict-pass multiple
Different-sample context
Official held-out results
- Gemini 3.7 Flash High30.44%
- Claude Opus 5 Max26.94%
- GPT-5.6 Terra Max21.00%
- GPT-5.6 Sol Max19.63%
Reference values come from the supplied official held-out comparison. They do not use the same sample as this 100-case frozen run, so they provide scale context only, not a model ranking.