Spaces:
Running
Running
win-speed bonus: record + reward how fast a model wins
Browse filesPer user: a faster win is worth a bonus and must be recorded/compared.
- scoring.py: _win_budget (tightest within_ticks in the win tree, else
max_ticks); ScoreCard gains win_tick/win_turns/win_budget/speed/
composite_base (in score.json + manifest). On a WIN, composite +=
SPEED_BONUS*speed (0.05 cap, speed=1-win_tick/budget) — orders fast
vs slow wins, never lifts a loss above a win or overrides correctness.
- run_eval._agg: win_speed_mean / win_turns_mean (over wins only).
- leaderboard: records win_speed/win_turns; composite/win_rate ties
now broken by faster wins.
- SCENARIO_QUALITY: #2 closer-look note + win-speed section.
+5 scoring/budget tests; suite 362 passed.
This view is limited to 50 files because it contains too many changes. See raw diff
- SCENARIO_QUALITY.md +48 -0
- eval_stats.json +27 -27
- openra_bench/leaderboard.py +8 -1
- openra_bench/run_eval.py +10 -0
- openra_bench/scoring.py +60 -2
- playback/run1/20260519-162229__qwen_qwen3.6-flash/_journal.jsonl +1 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/manifest.json +33 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/messages.json +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn001.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn002.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn003.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn004.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn005.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn006.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn007.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn008.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn009.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn010.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn011.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn012.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn013.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn014.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn015.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn016.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn017.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn018.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn019.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn020.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn021.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn022.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn023.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn024.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn025.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn026.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn027.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn028.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn029.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn030.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn031.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn032.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn033.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn034.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn035.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn036.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn037.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn038.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn039.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn040.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn041.png +0 -0
- playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn042.png +0 -0
SCENARIO_QUALITY.md
CHANGED
|
@@ -430,3 +430,51 @@ failure attributes to a single capability):
|
|
| 430 |
default `exact`) is a reusable knob: any pack can opt into
|
| 431 |
coordinate-blind objectives. The win predicate always evaluates
|
| 432 |
against the real engine coordinates regardless of disclosure.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 430 |
default `exact`) is a reusable knob: any pack can opt into
|
| 431 |
coordinate-blind objectives. The win predicate always evaluates
|
| 432 |
against the real engine coordinates regardless of disclosure.
|
| 433 |
+
|
| 434 |
+
## Per-scenario closer look — #2 action-sequenced-execution
|
| 435 |
+
|
| 436 |
+
Original defect: the pack claimed to test ordered, non-stalling route
|
| 437 |
+
execution but the win condition referenced only the *final* region +
|
| 438 |
+
a tick band — a beeline-to-final that skipped every waypoint and
|
| 439 |
+
idled to the gate won (verified: model arrived tick 1353, idled ~900
|
| 440 |
+
ticks, won). Not model compensation: the scenario didn't enforce what
|
| 441 |
+
it advertised.
|
| 442 |
+
|
| 443 |
+
Fix — new reusable **stateful** predicate `waypoint_sequence`: latches
|
| 444 |
+
ordered region-visit progress (monotonic) on a per-episode
|
| 445 |
+
`signals.seq_progress` scratch. Wk+1 only counts after Wk; skip /
|
| 446 |
+
out-of-order / idle ⇒ never satisfied. Coords-aware phrase
|
| 447 |
+
(`exact|relative`). Difficulty axis (one new controlled variable per
|
| 448 |
+
tier):
|
| 449 |
+
|
| 450 |
+
- **easy** — ONE ordered 3-waypoint route, clear lanes, generous
|
| 451 |
+
budget. Run: WIN, W1→W2→W3 in order, tick 1893 < 2400.
|
| 452 |
+
- **medium** — TWO ordered routes that *cross*, BOTH required ⇒ must
|
| 453 |
+
split into two parallel columns (serial overruns); attrition cap.
|
| 454 |
+
Run: LOSS (valid) — model split correctly, finished route S in
|
| 455 |
+
order, lost column discipline on N and timed out (the exact
|
| 456 |
+
re-deliberation failure the pack targets).
|
| 457 |
+
- **hard** — TWO 4-waypoint serpentine routes on the larger
|
| 458 |
+
`scout-arena` (176×80, `scripts/build_scout_arena_map.py`), given
|
| 459 |
+
ONLY by relative *search bands* (no coords); fogged markers revealed
|
| 460 |
+
by the `enemy_building_spotted` interrupt; seed-varied 2 spawn
|
| 461 |
+
points. Run: LOSS (valid) — machinery all verified working
|
| 462 |
+
(interrupts fired on the big map); model drove to literal corners
|
| 463 |
+
instead of scanning the bands, ignored the revealed positions,
|
| 464 |
+
attrition-bust. Solvable in principle; not compensated.
|
| 465 |
+
|
| 466 |
+
All tick budgets aligned to ~90 ticks/turn so timeout / attrition /
|
| 467 |
+
wipe are reachable **losses** within `max_turns` (no draw degeneracy).
|
| 468 |
+
Footgun recorded: a `base_map` placed at the Level level is silently
|
| 469 |
+
ignored — it must go inside `overrides:` (guard test added).
|
| 470 |
+
|
| 471 |
+
## Win-speed bonus (cross-cutting scoring)
|
| 472 |
+
|
| 473 |
+
Every episode records `win_tick` / `win_turns` / `win_budget` /
|
| 474 |
+
`speed` / `composite_base` (score.json + manifest). On a **win** the
|
| 475 |
+
composite gets `+SPEED_BONUS·speed` (`SPEED_BONUS=0.05`,
|
| 476 |
+
`speed = 1 − win_tick/budget`, budget = tightest `within_ticks` in the
|
| 477 |
+
win tree else `max_ticks`). Bounded so a fast win ranks above a slow
|
| 478 |
+
win but never lifts a loss above a win or overrides correctness.
|
| 479 |
+
Leaderboard carries `win_speed`/`win_turns` and breaks composite/
|
| 480 |
+
win-rate ties by faster wins.
|
eval_stats.json
CHANGED
|
@@ -1,25 +1,25 @@
|
|
| 1 |
{
|
| 2 |
-
"run_id": "20260519-
|
| 3 |
"model": "qwen/qwen3.6-flash",
|
| 4 |
"truncated": false,
|
| 5 |
"resumed": 0,
|
| 6 |
"cost": {
|
| 7 |
-
"calls":
|
| 8 |
-
"prompt_tokens":
|
| 9 |
-
"completion_tokens":
|
| 10 |
"usd": 0.0,
|
| 11 |
"max_usd": 0.0
|
| 12 |
},
|
| 13 |
"summary": {
|
| 14 |
-
"action-sequenced-execution:
|
| 15 |
"n": 1,
|
| 16 |
"win_rate": 0.0,
|
| 17 |
-
"composite_mean": 0.
|
| 18 |
"composite_std": 0.0,
|
| 19 |
-
"perception_mean": 0.
|
| 20 |
-
"reasoning_mean": 0.
|
| 21 |
"action_mean": 1.0,
|
| 22 |
-
"objective_mean": 0.
|
| 23 |
"weakest_link_hist": {
|
| 24 |
"reasoning": 1
|
| 25 |
}
|
|
@@ -28,44 +28,44 @@
|
|
| 28 |
"overall": {
|
| 29 |
"n": 1,
|
| 30 |
"win_rate": 0.0,
|
| 31 |
-
"composite_mean": 0.
|
| 32 |
"composite_std": 0.0,
|
| 33 |
-
"perception_mean": 0.
|
| 34 |
-
"reasoning_mean": 0.
|
| 35 |
"action_mean": 1.0,
|
| 36 |
-
"objective_mean": 0.
|
| 37 |
"weakest_link_hist": {
|
| 38 |
"reasoning": 1
|
| 39 |
}
|
| 40 |
},
|
| 41 |
"reward_vector_mean": {
|
| 42 |
"economy": 0.5,
|
| 43 |
-
"military": 0.
|
| 44 |
-
"territory": 0.
|
| 45 |
-
"scouting": 0.
|
| 46 |
-
"objective": 0.
|
| 47 |
},
|
| 48 |
"episodes": [
|
| 49 |
{
|
| 50 |
-
"cell": "action-sequenced-execution:
|
| 51 |
"capability": "action",
|
| 52 |
"split": "public",
|
| 53 |
"seed": 1,
|
| 54 |
"outcome": "loss",
|
| 55 |
-
"composite": 0.
|
| 56 |
-
"perception": 0.
|
| 57 |
-
"reasoning": 0.
|
| 58 |
"action": 1.0,
|
| 59 |
"weakest_link": "reasoning",
|
| 60 |
-
"objective_progress": 0.
|
| 61 |
"reward_vector": {
|
| 62 |
"economy": 0.5,
|
| 63 |
-
"military": 0.
|
| 64 |
-
"territory": 0.
|
| 65 |
-
"scouting": 0.
|
| 66 |
-
"objective": 0.
|
| 67 |
},
|
| 68 |
-
"turns":
|
| 69 |
"notes": [
|
| 70 |
"objective not met (loss); weakest link: reasoning"
|
| 71 |
]
|
|
|
|
| 1 |
{
|
| 2 |
+
"run_id": "20260519-162229",
|
| 3 |
"model": "qwen/qwen3.6-flash",
|
| 4 |
"truncated": false,
|
| 5 |
"resumed": 0,
|
| 6 |
"cost": {
|
| 7 |
+
"calls": 41,
|
| 8 |
+
"prompt_tokens": 332060,
|
| 9 |
+
"completion_tokens": 43651,
|
| 10 |
"usd": 0.0,
|
| 11 |
"max_usd": 0.0
|
| 12 |
},
|
| 13 |
"summary": {
|
| 14 |
+
"action-sequenced-execution:hard": {
|
| 15 |
"n": 1,
|
| 16 |
"win_rate": 0.0,
|
| 17 |
+
"composite_mean": 0.1773,
|
| 18 |
"composite_std": 0.0,
|
| 19 |
+
"perception_mean": 0.6844,
|
| 20 |
+
"reasoning_mean": 0.6737,
|
| 21 |
"action_mean": 1.0,
|
| 22 |
+
"objective_mean": 0.375,
|
| 23 |
"weakest_link_hist": {
|
| 24 |
"reasoning": 1
|
| 25 |
}
|
|
|
|
| 28 |
"overall": {
|
| 29 |
"n": 1,
|
| 30 |
"win_rate": 0.0,
|
| 31 |
+
"composite_mean": 0.1773,
|
| 32 |
"composite_std": 0.0,
|
| 33 |
+
"perception_mean": 0.6844,
|
| 34 |
+
"reasoning_mean": 0.6737,
|
| 35 |
"action_mean": 1.0,
|
| 36 |
+
"objective_mean": 0.375,
|
| 37 |
"weakest_link_hist": {
|
| 38 |
"reasoning": 1
|
| 39 |
}
|
| 40 |
},
|
| 41 |
"reward_vector_mean": {
|
| 42 |
"economy": 0.5,
|
| 43 |
+
"military": 0.0,
|
| 44 |
+
"territory": 0.5491,
|
| 45 |
+
"scouting": 0.6,
|
| 46 |
+
"objective": 0.375
|
| 47 |
},
|
| 48 |
"episodes": [
|
| 49 |
{
|
| 50 |
+
"cell": "action-sequenced-execution:hard",
|
| 51 |
"capability": "action",
|
| 52 |
"split": "public",
|
| 53 |
"seed": 1,
|
| 54 |
"outcome": "loss",
|
| 55 |
+
"composite": 0.1773,
|
| 56 |
+
"perception": 0.6844,
|
| 57 |
+
"reasoning": 0.6737,
|
| 58 |
"action": 1.0,
|
| 59 |
"weakest_link": "reasoning",
|
| 60 |
+
"objective_progress": 0.375,
|
| 61 |
"reward_vector": {
|
| 62 |
"economy": 0.5,
|
| 63 |
+
"military": 0.0,
|
| 64 |
+
"territory": 0.5491,
|
| 65 |
+
"scouting": 0.6,
|
| 66 |
+
"objective": 0.375
|
| 67 |
},
|
| 68 |
+
"turns": 41,
|
| 69 |
"notes": [
|
| 70 |
"objective not met (loss); weakest link: reasoning"
|
| 71 |
]
|
openra_bench/leaderboard.py
CHANGED
|
@@ -68,6 +68,9 @@ def ingest_run(
|
|
| 68 |
# Continuous goal achievement (partial credit) + the cumulative
|
| 69 |
# reward-vector signature, comparable across models/runs.
|
| 70 |
"objective": overall.get("objective_mean", 0.0),
|
|
|
|
|
|
|
|
|
|
| 71 |
"reward_vector": stats.get("reward_vector_mean", {}),
|
| 72 |
# Adversarial 1v1 spotlight: mean ladder rating (0–3) + per-pack.
|
| 73 |
"adversarial_rating": stats.get("adversarial", {}).get(
|
|
@@ -125,7 +128,11 @@ def build_table(
|
|
| 125 |
best[m] = r
|
| 126 |
rows = sorted(
|
| 127 |
best.values(),
|
| 128 |
-
key=lambda r: (
|
|
|
|
|
|
|
|
|
|
|
|
|
| 129 |
)
|
| 130 |
for i, r in enumerate(rows, 1):
|
| 131 |
r["rank"] = i
|
|
|
|
| 68 |
# Continuous goal achievement (partial credit) + the cumulative
|
| 69 |
# reward-vector signature, comparable across models/runs.
|
| 70 |
"objective": overall.get("objective_mean", 0.0),
|
| 71 |
+
# Win speed (mean over wins): how decisively the model wins.
|
| 72 |
+
"win_speed": overall.get("win_speed_mean", 0.0),
|
| 73 |
+
"win_turns": overall.get("win_turns_mean", 0.0),
|
| 74 |
"reward_vector": stats.get("reward_vector_mean", {}),
|
| 75 |
# Adversarial 1v1 spotlight: mean ladder rating (0–3) + per-pack.
|
| 76 |
"adversarial_rating": stats.get("adversarial", {}).get(
|
|
|
|
| 128 |
best[m] = r
|
| 129 |
rows = sorted(
|
| 130 |
best.values(),
|
| 131 |
+
key=lambda r: (
|
| 132 |
+
-r["composite"], -r["win_rate"],
|
| 133 |
+
-r.get("win_speed", 0.0), # faster wins break ties
|
| 134 |
+
r["model"],
|
| 135 |
+
),
|
| 136 |
)
|
| 137 |
for i, r in enumerate(rows, 1):
|
| 138 |
r["rank"] = i
|
openra_bench/run_eval.py
CHANGED
|
@@ -85,6 +85,16 @@ def _agg(scores: list) -> dict:
|
|
| 85 |
"objective_mean": round(
|
| 86 |
statistics.fmean(s.dimensions.get("objective", 0.0) for s in scores), 4
|
| 87 |
),
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 88 |
"weakest_link_hist": dict(Counter(s.weakest_link for s in scores)),
|
| 89 |
}
|
| 90 |
|
|
|
|
| 85 |
"objective_mean": round(
|
| 86 |
statistics.fmean(s.dimensions.get("objective", 0.0) for s in scores), 4
|
| 87 |
),
|
| 88 |
+
# Win-speed: averaged over WINS only (0 when there are none) so
|
| 89 |
+
# it compares how decisively a model wins, not diluted by losses.
|
| 90 |
+
"win_speed_mean": round(
|
| 91 |
+
statistics.fmean([s.speed for s in scores if s.outcome == "win"]), 4
|
| 92 |
+
) if any(s.outcome == "win" for s in scores) else 0.0,
|
| 93 |
+
"win_turns_mean": round(
|
| 94 |
+
statistics.fmean(
|
| 95 |
+
[s.win_turns for s in scores if s.outcome == "win"]
|
| 96 |
+
), 2
|
| 97 |
+
) if any(s.outcome == "win" for s in scores) else 0.0,
|
| 98 |
"weakest_link_hist": dict(Counter(s.weakest_link for s in scores)),
|
| 99 |
}
|
| 100 |
|
openra_bench/scoring.py
CHANGED
|
@@ -43,7 +43,7 @@ def _clamp(x: float, lo: float = 0.0, hi: float = 1.0) -> float:
|
|
| 43 |
|
| 44 |
@dataclass
|
| 45 |
class ScoreCard:
|
| 46 |
-
composite: float # weighted scalar in [0,1]
|
| 47 |
outcome: str # win | draw | loss
|
| 48 |
perception: float # P/R/A link sub-scores, each in [0,1]
|
| 49 |
reasoning: float
|
|
@@ -52,11 +52,52 @@ class ScoreCard:
|
|
| 52 |
weights: dict # weights actually used (scenario or default)
|
| 53 |
weakest_link: str # "perception" | "reasoning" | "action"
|
| 54 |
notes: list
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
def to_dict(self) -> dict:
|
| 57 |
return asdict(self)
|
| 58 |
|
| 59 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 60 |
def _dimension_values(compiled: CompiledLevel, res: EpisodeResult) -> dict:
|
| 61 |
"""Map adapter signals -> Training reward dimensions, each in [0,1].
|
| 62 |
|
|
@@ -172,11 +213,23 @@ def _pra_diagnostics(compiled: CompiledLevel, res: EpisodeResult, dims: dict) ->
|
|
| 172 |
def score_episode(compiled: CompiledLevel, res: EpisodeResult) -> ScoreCard:
|
| 173 |
dims = _dimension_values(compiled, res)
|
| 174 |
weights = _weights(compiled)
|
| 175 |
-
|
| 176 |
pra = _pra_diagnostics(compiled, res, dims)
|
| 177 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 178 |
weakest = min(pra, key=pra.get)
|
| 179 |
notes: list[str] = []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 180 |
if res.outcome != "win":
|
| 181 |
notes.append(f"objective not met ({res.outcome}); weakest link: {weakest}")
|
| 182 |
if res.actions_issued and res.actions_warned / res.actions_issued > 0.25:
|
|
@@ -197,4 +250,9 @@ def score_episode(compiled: CompiledLevel, res: EpisodeResult) -> ScoreCard:
|
|
| 197 |
weights=weights,
|
| 198 |
weakest_link=weakest,
|
| 199 |
notes=notes,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 200 |
)
|
|
|
|
| 43 |
|
| 44 |
@dataclass
|
| 45 |
class ScoreCard:
|
| 46 |
+
composite: float # weighted scalar in [0,1] (incl. speed bonus)
|
| 47 |
outcome: str # win | draw | loss
|
| 48 |
perception: float # P/R/A link sub-scores, each in [0,1]
|
| 49 |
reasoning: float
|
|
|
|
| 52 |
weights: dict # weights actually used (scenario or default)
|
| 53 |
weakest_link: str # "perception" | "reasoning" | "action"
|
| 54 |
notes: list
|
| 55 |
+
# Win-speed bonus (recorded for every episode; only non-zero on a
|
| 56 |
+
# win). speed ∈ [0,1] = 1 − win_tick/budget (faster ⇒ higher); the
|
| 57 |
+
# composite gets at most SPEED_BONUS·speed added — enough to rank
|
| 58 |
+
# fast wins above slow wins, never enough to lift a loss above a
|
| 59 |
+
# win or override correctness.
|
| 60 |
+
win_tick: int = 0 # game tick the win fired (0 if not won)
|
| 61 |
+
win_turns: int = 0 # decision turns to the win (0 if not won)
|
| 62 |
+
win_budget: int = 0 # the tick budget judged against
|
| 63 |
+
speed: float = 0.0 # [0,1], 0 unless won
|
| 64 |
+
composite_base: float = 0.0 # composite before the speed bonus
|
| 65 |
|
| 66 |
def to_dict(self) -> dict:
|
| 67 |
return asdict(self)
|
| 68 |
|
| 69 |
|
| 70 |
+
# Max additive speed bonus on the composite (wins only). Small by
|
| 71 |
+
# design: orders fast vs slow wins without dominating correctness.
|
| 72 |
+
SPEED_BONUS = 0.05
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
def _win_budget(compiled: CompiledLevel) -> int:
|
| 76 |
+
"""Tick budget a win is judged 'fast' against: the tightest
|
| 77 |
+
`within_ticks` in the win tree, else the scenario max_ticks."""
|
| 78 |
+
best: list[int] = []
|
| 79 |
+
|
| 80 |
+
def walk(node):
|
| 81 |
+
if node is None:
|
| 82 |
+
return
|
| 83 |
+
d = node if isinstance(node, dict) else dict(
|
| 84 |
+
getattr(node, "__pydantic_extra__", {}) or {}
|
| 85 |
+
)
|
| 86 |
+
for k, v in d.items():
|
| 87 |
+
if k in ("all_of", "any_of"):
|
| 88 |
+
for c in v:
|
| 89 |
+
walk(c)
|
| 90 |
+
elif k == "not":
|
| 91 |
+
walk(v)
|
| 92 |
+
elif k == "within_ticks":
|
| 93 |
+
best.append(int(v))
|
| 94 |
+
|
| 95 |
+
walk(compiled.win_condition)
|
| 96 |
+
if best:
|
| 97 |
+
return max(1, min(best))
|
| 98 |
+
return max(1, compiled.scenario.termination.max_ticks)
|
| 99 |
+
|
| 100 |
+
|
| 101 |
def _dimension_values(compiled: CompiledLevel, res: EpisodeResult) -> dict:
|
| 102 |
"""Map adapter signals -> Training reward dimensions, each in [0,1].
|
| 103 |
|
|
|
|
| 213 |
def score_episode(compiled: CompiledLevel, res: EpisodeResult) -> ScoreCard:
|
| 214 |
dims = _dimension_values(compiled, res)
|
| 215 |
weights = _weights(compiled)
|
| 216 |
+
composite_base = _composite(dims, weights)
|
| 217 |
pra = _pra_diagnostics(compiled, res, dims)
|
| 218 |
|
| 219 |
+
won = res.outcome == "win"
|
| 220 |
+
win_tick = int(res.signals.game_tick) if won else 0
|
| 221 |
+
win_turns = int(res.turns) if won else 0
|
| 222 |
+
budget = _win_budget(compiled)
|
| 223 |
+
speed = _clamp(1.0 - win_tick / budget) if won and win_tick > 0 else 0.0
|
| 224 |
+
composite = _clamp(composite_base + SPEED_BONUS * speed) if won else composite_base
|
| 225 |
+
|
| 226 |
weakest = min(pra, key=pra.get)
|
| 227 |
notes: list[str] = []
|
| 228 |
+
if won:
|
| 229 |
+
notes.append(
|
| 230 |
+
f"won in {win_turns} turns / tick {win_tick} of {budget} "
|
| 231 |
+
f"(speed {speed:.2f}, +{SPEED_BONUS * speed:.3f} bonus)"
|
| 232 |
+
)
|
| 233 |
if res.outcome != "win":
|
| 234 |
notes.append(f"objective not met ({res.outcome}); weakest link: {weakest}")
|
| 235 |
if res.actions_issued and res.actions_warned / res.actions_issued > 0.25:
|
|
|
|
| 250 |
weights=weights,
|
| 251 |
weakest_link=weakest,
|
| 252 |
notes=notes,
|
| 253 |
+
win_tick=win_tick,
|
| 254 |
+
win_turns=win_turns,
|
| 255 |
+
win_budget=budget,
|
| 256 |
+
speed=round(speed, 4),
|
| 257 |
+
composite_base=round(composite_base, 4),
|
| 258 |
)
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/_journal.jsonl
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
{"cell": "action-sequenced-execution:hard", "capability": "action", "split": "public", "seed": 1, "outcome": "loss", "composite": 0.1773, "perception": 0.6844, "reasoning": 0.6737, "action": 1.0, "weakest_link": "reasoning", "objective_progress": 0.375, "reward_vector": {"economy": 0.5, "military": 0.0, "territory": 0.5491, "scouting": 0.6, "objective": 0.375}, "turns": 41, "notes": ["objective not met (loss); weakest link: reasoning"], "_key": "action-sequenced-execution|hard|public|1"}
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/manifest.json
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"scenario": "action-sequenced-execution:hard",
|
| 3 |
+
"pack_id": "action-sequenced-execution",
|
| 4 |
+
"level": "hard",
|
| 5 |
+
"capability": "action",
|
| 6 |
+
"run_id": "20260519-162229",
|
| 7 |
+
"model": "qwen/qwen3.6-flash",
|
| 8 |
+
"seed": 1,
|
| 9 |
+
"outcome": "loss",
|
| 10 |
+
"turns": 41,
|
| 11 |
+
"max_turns": 70,
|
| 12 |
+
"actions_issued": 78,
|
| 13 |
+
"actions_warned": 0,
|
| 14 |
+
"agent_stats": {
|
| 15 |
+
"turns": 41,
|
| 16 |
+
"tool_calls": 78,
|
| 17 |
+
"empty_replies": 0
|
| 18 |
+
},
|
| 19 |
+
"objective_progress": 0.375,
|
| 20 |
+
"reward_vector": {
|
| 21 |
+
"economy": 0.5,
|
| 22 |
+
"military": 0.0,
|
| 23 |
+
"territory": 0.5491,
|
| 24 |
+
"scouting": 0.6,
|
| 25 |
+
"objective": 0.375
|
| 26 |
+
},
|
| 27 |
+
"signals": {
|
| 28 |
+
"economy_value": 5000,
|
| 29 |
+
"explored_percent": 54.91,
|
| 30 |
+
"units_killed": 0,
|
| 31 |
+
"units_lost": 2
|
| 32 |
+
}
|
| 33 |
+
}
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/messages.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn001.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn002.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn003.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn004.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn005.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn006.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn007.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn008.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn009.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn010.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn011.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn012.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn013.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn014.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn015.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn016.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn017.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn018.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn019.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn020.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn021.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn022.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn023.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn024.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn025.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn026.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn027.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn028.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn029.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn030.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn031.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn032.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn033.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn034.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn035.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn036.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn037.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn038.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn039.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn040.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn041.png
ADDED
|
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn042.png
ADDED
|