Best pass@1
74.2%Orion-3.1 · +8.4 ptsObserve
Runs & Analytics
Compare model snapshots on frozen environments and inspect the trajectories behind every result.
Reproducibility
99.8%2 inconsistent of 1,248 runsMedian episode
04:38p90 · 09:12Cost / success
$0.71−12% vs. prior checkpointModel comparison
Pass@1 across 24 frozen environments
Capability matrix
Success rate by model and capability
OrionNovaAtlasBaseline
Code repair82%68%56%34%
Reasoning71%77%43%28%
Tool use69%52%45%24%
Robustness76%49%51%31%
LowerHigher
TOP FAILURE CLUSTERIncomplete state recovery
18 episodes · mostly concurrent code tasks
LARGEST REGRESSIONLong-horizon tool use
−9.2 pts vs. Orion-3.0
LARGEST GAINAdversarial parsing
+14.8 pts vs. prior checkpoint
Recent episodes
Exact model, environment revision, prompt, and result for every run
| Run | Model snapshot | Environment | Result | Reward | Duration | Cost | Started |
|---|---|---|---|---|---|---|---|
run_10842 | Orion-3.1 | Atomic cache recovery | Passed | 1.00 | 04:18 | $0.42 | 4 min ago |
run_10841 | Atlas-Coder-70B | Atomic cache recovery | Failed | 0.00 | 06:00 | $0.31 | 12 min ago |
run_10840 | Orion-3.1 | Incident triage desk | Passed | 0.92 | 08:47 | $0.88 | 21 min ago |
run_10839 | Nova-Reasoner | Bounded permutation orbits | Failed | 0.00 | 02:11 | $0.19 | 34 min ago |
run_10838 | Atlas-Coder-70B | Streaming ledger parser | Passed | 1.00 | 05:26 | $0.37 | 51 min ago |