yxc20098 commited on
Commit
35d7d47
·
1 Parent(s): cdd813f

win-speed bonus: record + reward how fast a model wins

Browse files

Per user: a faster win is worth a bonus and must be recorded/compared.
- scoring.py: _win_budget (tightest within_ticks in the win tree, else
max_ticks); ScoreCard gains win_tick/win_turns/win_budget/speed/
composite_base (in score.json + manifest). On a WIN, composite +=
SPEED_BONUS*speed (0.05 cap, speed=1-win_tick/budget) — orders fast
vs slow wins, never lifts a loss above a win or overrides correctness.
- run_eval._agg: win_speed_mean / win_turns_mean (over wins only).
- leaderboard: records win_speed/win_turns; composite/win_rate ties
now broken by faster wins.
- SCENARIO_QUALITY: #2 closer-look note + win-speed section.
+5 scoring/budget tests; suite 362 passed.

This view is limited to 50 files because it contains too many changes.   See raw diff
Files changed (50) hide show
  1. SCENARIO_QUALITY.md +48 -0
  2. eval_stats.json +27 -27
  3. openra_bench/leaderboard.py +8 -1
  4. openra_bench/run_eval.py +10 -0
  5. openra_bench/scoring.py +60 -2
  6. playback/run1/20260519-162229__qwen_qwen3.6-flash/_journal.jsonl +1 -0
  7. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/manifest.json +33 -0
  8. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/messages.json +0 -0
  9. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn001.png +0 -0
  10. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn002.png +0 -0
  11. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn003.png +0 -0
  12. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn004.png +0 -0
  13. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn005.png +0 -0
  14. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn006.png +0 -0
  15. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn007.png +0 -0
  16. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn008.png +0 -0
  17. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn009.png +0 -0
  18. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn010.png +0 -0
  19. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn011.png +0 -0
  20. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn012.png +0 -0
  21. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn013.png +0 -0
  22. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn014.png +0 -0
  23. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn015.png +0 -0
  24. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn016.png +0 -0
  25. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn017.png +0 -0
  26. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn018.png +0 -0
  27. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn019.png +0 -0
  28. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn020.png +0 -0
  29. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn021.png +0 -0
  30. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn022.png +0 -0
  31. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn023.png +0 -0
  32. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn024.png +0 -0
  33. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn025.png +0 -0
  34. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn026.png +0 -0
  35. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn027.png +0 -0
  36. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn028.png +0 -0
  37. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn029.png +0 -0
  38. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn030.png +0 -0
  39. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn031.png +0 -0
  40. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn032.png +0 -0
  41. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn033.png +0 -0
  42. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn034.png +0 -0
  43. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn035.png +0 -0
  44. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn036.png +0 -0
  45. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn037.png +0 -0
  46. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn038.png +0 -0
  47. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn039.png +0 -0
  48. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn040.png +0 -0
  49. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn041.png +0 -0
  50. playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn042.png +0 -0
SCENARIO_QUALITY.md CHANGED
@@ -430,3 +430,51 @@ failure attributes to a single capability):
430
  default `exact`) is a reusable knob: any pack can opt into
431
  coordinate-blind objectives. The win predicate always evaluates
432
  against the real engine coordinates regardless of disclosure.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
430
  default `exact`) is a reusable knob: any pack can opt into
431
  coordinate-blind objectives. The win predicate always evaluates
432
  against the real engine coordinates regardless of disclosure.
433
+
434
+ ## Per-scenario closer look — #2 action-sequenced-execution
435
+
436
+ Original defect: the pack claimed to test ordered, non-stalling route
437
+ execution but the win condition referenced only the *final* region +
438
+ a tick band — a beeline-to-final that skipped every waypoint and
439
+ idled to the gate won (verified: model arrived tick 1353, idled ~900
440
+ ticks, won). Not model compensation: the scenario didn't enforce what
441
+ it advertised.
442
+
443
+ Fix — new reusable **stateful** predicate `waypoint_sequence`: latches
444
+ ordered region-visit progress (monotonic) on a per-episode
445
+ `signals.seq_progress` scratch. Wk+1 only counts after Wk; skip /
446
+ out-of-order / idle ⇒ never satisfied. Coords-aware phrase
447
+ (`exact|relative`). Difficulty axis (one new controlled variable per
448
+ tier):
449
+
450
+ - **easy** — ONE ordered 3-waypoint route, clear lanes, generous
451
+ budget. Run: WIN, W1→W2→W3 in order, tick 1893 < 2400.
452
+ - **medium** — TWO ordered routes that *cross*, BOTH required ⇒ must
453
+ split into two parallel columns (serial overruns); attrition cap.
454
+ Run: LOSS (valid) — model split correctly, finished route S in
455
+ order, lost column discipline on N and timed out (the exact
456
+ re-deliberation failure the pack targets).
457
+ - **hard** — TWO 4-waypoint serpentine routes on the larger
458
+ `scout-arena` (176×80, `scripts/build_scout_arena_map.py`), given
459
+ ONLY by relative *search bands* (no coords); fogged markers revealed
460
+ by the `enemy_building_spotted` interrupt; seed-varied 2 spawn
461
+ points. Run: LOSS (valid) — machinery all verified working
462
+ (interrupts fired on the big map); model drove to literal corners
463
+ instead of scanning the bands, ignored the revealed positions,
464
+ attrition-bust. Solvable in principle; not compensated.
465
+
466
+ All tick budgets aligned to ~90 ticks/turn so timeout / attrition /
467
+ wipe are reachable **losses** within `max_turns` (no draw degeneracy).
468
+ Footgun recorded: a `base_map` placed at the Level level is silently
469
+ ignored — it must go inside `overrides:` (guard test added).
470
+
471
+ ## Win-speed bonus (cross-cutting scoring)
472
+
473
+ Every episode records `win_tick` / `win_turns` / `win_budget` /
474
+ `speed` / `composite_base` (score.json + manifest). On a **win** the
475
+ composite gets `+SPEED_BONUS·speed` (`SPEED_BONUS=0.05`,
476
+ `speed = 1 − win_tick/budget`, budget = tightest `within_ticks` in the
477
+ win tree else `max_ticks`). Bounded so a fast win ranks above a slow
478
+ win but never lifts a loss above a win or overrides correctness.
479
+ Leaderboard carries `win_speed`/`win_turns` and breaks composite/
480
+ win-rate ties by faster wins.
eval_stats.json CHANGED
@@ -1,25 +1,25 @@
1
  {
2
- "run_id": "20260519-155206",
3
  "model": "qwen/qwen3.6-flash",
4
  "truncated": false,
5
  "resumed": 0,
6
  "cost": {
7
- "calls": 34,
8
- "prompt_tokens": 224195,
9
- "completion_tokens": 34787,
10
  "usd": 0.0,
11
  "max_usd": 0.0
12
  },
13
  "summary": {
14
- "action-sequenced-execution:medium": {
15
  "n": 1,
16
  "win_rate": 0.0,
17
- "composite_mean": 0.3749,
18
  "composite_std": 0.0,
19
- "perception_mean": 0.868,
20
- "reasoning_mean": 0.8663,
21
  "action_mean": 1.0,
22
- "objective_mean": 0.7449,
23
  "weakest_link_hist": {
24
  "reasoning": 1
25
  }
@@ -28,44 +28,44 @@
28
  "overall": {
29
  "n": 1,
30
  "win_rate": 0.0,
31
- "composite_mean": 0.3749,
32
  "composite_std": 0.0,
33
- "perception_mean": 0.868,
34
- "reasoning_mean": 0.8663,
35
  "action_mean": 1.0,
36
- "objective_mean": 0.7449,
37
  "weakest_link_hist": {
38
  "reasoning": 1
39
  }
40
  },
41
  "reward_vector_mean": {
42
  "economy": 0.5,
43
- "military": 0.4,
44
- "territory": 0.8114,
45
- "scouting": 0.7,
46
- "objective": 0.7449
47
  },
48
  "episodes": [
49
  {
50
- "cell": "action-sequenced-execution:medium",
51
  "capability": "action",
52
  "split": "public",
53
  "seed": 1,
54
  "outcome": "loss",
55
- "composite": 0.3749,
56
- "perception": 0.868,
57
- "reasoning": 0.8663,
58
  "action": 1.0,
59
  "weakest_link": "reasoning",
60
- "objective_progress": 0.7449,
61
  "reward_vector": {
62
  "economy": 0.5,
63
- "military": 0.4,
64
- "territory": 0.8114,
65
- "scouting": 0.7,
66
- "objective": 0.7449
67
  },
68
- "turns": 34,
69
  "notes": [
70
  "objective not met (loss); weakest link: reasoning"
71
  ]
 
1
  {
2
+ "run_id": "20260519-162229",
3
  "model": "qwen/qwen3.6-flash",
4
  "truncated": false,
5
  "resumed": 0,
6
  "cost": {
7
+ "calls": 41,
8
+ "prompt_tokens": 332060,
9
+ "completion_tokens": 43651,
10
  "usd": 0.0,
11
  "max_usd": 0.0
12
  },
13
  "summary": {
14
+ "action-sequenced-execution:hard": {
15
  "n": 1,
16
  "win_rate": 0.0,
17
+ "composite_mean": 0.1773,
18
  "composite_std": 0.0,
19
+ "perception_mean": 0.6844,
20
+ "reasoning_mean": 0.6737,
21
  "action_mean": 1.0,
22
+ "objective_mean": 0.375,
23
  "weakest_link_hist": {
24
  "reasoning": 1
25
  }
 
28
  "overall": {
29
  "n": 1,
30
  "win_rate": 0.0,
31
+ "composite_mean": 0.1773,
32
  "composite_std": 0.0,
33
+ "perception_mean": 0.6844,
34
+ "reasoning_mean": 0.6737,
35
  "action_mean": 1.0,
36
+ "objective_mean": 0.375,
37
  "weakest_link_hist": {
38
  "reasoning": 1
39
  }
40
  },
41
  "reward_vector_mean": {
42
  "economy": 0.5,
43
+ "military": 0.0,
44
+ "territory": 0.5491,
45
+ "scouting": 0.6,
46
+ "objective": 0.375
47
  },
48
  "episodes": [
49
  {
50
+ "cell": "action-sequenced-execution:hard",
51
  "capability": "action",
52
  "split": "public",
53
  "seed": 1,
54
  "outcome": "loss",
55
+ "composite": 0.1773,
56
+ "perception": 0.6844,
57
+ "reasoning": 0.6737,
58
  "action": 1.0,
59
  "weakest_link": "reasoning",
60
+ "objective_progress": 0.375,
61
  "reward_vector": {
62
  "economy": 0.5,
63
+ "military": 0.0,
64
+ "territory": 0.5491,
65
+ "scouting": 0.6,
66
+ "objective": 0.375
67
  },
68
+ "turns": 41,
69
  "notes": [
70
  "objective not met (loss); weakest link: reasoning"
71
  ]
openra_bench/leaderboard.py CHANGED
@@ -68,6 +68,9 @@ def ingest_run(
68
  # Continuous goal achievement (partial credit) + the cumulative
69
  # reward-vector signature, comparable across models/runs.
70
  "objective": overall.get("objective_mean", 0.0),
 
 
 
71
  "reward_vector": stats.get("reward_vector_mean", {}),
72
  # Adversarial 1v1 spotlight: mean ladder rating (0–3) + per-pack.
73
  "adversarial_rating": stats.get("adversarial", {}).get(
@@ -125,7 +128,11 @@ def build_table(
125
  best[m] = r
126
  rows = sorted(
127
  best.values(),
128
- key=lambda r: (-r["composite"], -r["win_rate"], r["model"]),
 
 
 
 
129
  )
130
  for i, r in enumerate(rows, 1):
131
  r["rank"] = i
 
68
  # Continuous goal achievement (partial credit) + the cumulative
69
  # reward-vector signature, comparable across models/runs.
70
  "objective": overall.get("objective_mean", 0.0),
71
+ # Win speed (mean over wins): how decisively the model wins.
72
+ "win_speed": overall.get("win_speed_mean", 0.0),
73
+ "win_turns": overall.get("win_turns_mean", 0.0),
74
  "reward_vector": stats.get("reward_vector_mean", {}),
75
  # Adversarial 1v1 spotlight: mean ladder rating (0–3) + per-pack.
76
  "adversarial_rating": stats.get("adversarial", {}).get(
 
128
  best[m] = r
129
  rows = sorted(
130
  best.values(),
131
+ key=lambda r: (
132
+ -r["composite"], -r["win_rate"],
133
+ -r.get("win_speed", 0.0), # faster wins break ties
134
+ r["model"],
135
+ ),
136
  )
137
  for i, r in enumerate(rows, 1):
138
  r["rank"] = i
openra_bench/run_eval.py CHANGED
@@ -85,6 +85,16 @@ def _agg(scores: list) -> dict:
85
  "objective_mean": round(
86
  statistics.fmean(s.dimensions.get("objective", 0.0) for s in scores), 4
87
  ),
 
 
 
 
 
 
 
 
 
 
88
  "weakest_link_hist": dict(Counter(s.weakest_link for s in scores)),
89
  }
90
 
 
85
  "objective_mean": round(
86
  statistics.fmean(s.dimensions.get("objective", 0.0) for s in scores), 4
87
  ),
88
+ # Win-speed: averaged over WINS only (0 when there are none) so
89
+ # it compares how decisively a model wins, not diluted by losses.
90
+ "win_speed_mean": round(
91
+ statistics.fmean([s.speed for s in scores if s.outcome == "win"]), 4
92
+ ) if any(s.outcome == "win" for s in scores) else 0.0,
93
+ "win_turns_mean": round(
94
+ statistics.fmean(
95
+ [s.win_turns for s in scores if s.outcome == "win"]
96
+ ), 2
97
+ ) if any(s.outcome == "win" for s in scores) else 0.0,
98
  "weakest_link_hist": dict(Counter(s.weakest_link for s in scores)),
99
  }
100
 
openra_bench/scoring.py CHANGED
@@ -43,7 +43,7 @@ def _clamp(x: float, lo: float = 0.0, hi: float = 1.0) -> float:
43
 
44
  @dataclass
45
  class ScoreCard:
46
- composite: float # weighted scalar in [0,1]
47
  outcome: str # win | draw | loss
48
  perception: float # P/R/A link sub-scores, each in [0,1]
49
  reasoning: float
@@ -52,11 +52,52 @@ class ScoreCard:
52
  weights: dict # weights actually used (scenario or default)
53
  weakest_link: str # "perception" | "reasoning" | "action"
54
  notes: list
 
 
 
 
 
 
 
 
 
 
55
 
56
  def to_dict(self) -> dict:
57
  return asdict(self)
58
 
59
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
60
  def _dimension_values(compiled: CompiledLevel, res: EpisodeResult) -> dict:
61
  """Map adapter signals -> Training reward dimensions, each in [0,1].
62
 
@@ -172,11 +213,23 @@ def _pra_diagnostics(compiled: CompiledLevel, res: EpisodeResult, dims: dict) ->
172
  def score_episode(compiled: CompiledLevel, res: EpisodeResult) -> ScoreCard:
173
  dims = _dimension_values(compiled, res)
174
  weights = _weights(compiled)
175
- composite = _composite(dims, weights)
176
  pra = _pra_diagnostics(compiled, res, dims)
177
 
 
 
 
 
 
 
 
178
  weakest = min(pra, key=pra.get)
179
  notes: list[str] = []
 
 
 
 
 
180
  if res.outcome != "win":
181
  notes.append(f"objective not met ({res.outcome}); weakest link: {weakest}")
182
  if res.actions_issued and res.actions_warned / res.actions_issued > 0.25:
@@ -197,4 +250,9 @@ def score_episode(compiled: CompiledLevel, res: EpisodeResult) -> ScoreCard:
197
  weights=weights,
198
  weakest_link=weakest,
199
  notes=notes,
 
 
 
 
 
200
  )
 
43
 
44
  @dataclass
45
  class ScoreCard:
46
+ composite: float # weighted scalar in [0,1] (incl. speed bonus)
47
  outcome: str # win | draw | loss
48
  perception: float # P/R/A link sub-scores, each in [0,1]
49
  reasoning: float
 
52
  weights: dict # weights actually used (scenario or default)
53
  weakest_link: str # "perception" | "reasoning" | "action"
54
  notes: list
55
+ # Win-speed bonus (recorded for every episode; only non-zero on a
56
+ # win). speed ∈ [0,1] = 1 − win_tick/budget (faster ⇒ higher); the
57
+ # composite gets at most SPEED_BONUS·speed added — enough to rank
58
+ # fast wins above slow wins, never enough to lift a loss above a
59
+ # win or override correctness.
60
+ win_tick: int = 0 # game tick the win fired (0 if not won)
61
+ win_turns: int = 0 # decision turns to the win (0 if not won)
62
+ win_budget: int = 0 # the tick budget judged against
63
+ speed: float = 0.0 # [0,1], 0 unless won
64
+ composite_base: float = 0.0 # composite before the speed bonus
65
 
66
  def to_dict(self) -> dict:
67
  return asdict(self)
68
 
69
 
70
+ # Max additive speed bonus on the composite (wins only). Small by
71
+ # design: orders fast vs slow wins without dominating correctness.
72
+ SPEED_BONUS = 0.05
73
+
74
+
75
+ def _win_budget(compiled: CompiledLevel) -> int:
76
+ """Tick budget a win is judged 'fast' against: the tightest
77
+ `within_ticks` in the win tree, else the scenario max_ticks."""
78
+ best: list[int] = []
79
+
80
+ def walk(node):
81
+ if node is None:
82
+ return
83
+ d = node if isinstance(node, dict) else dict(
84
+ getattr(node, "__pydantic_extra__", {}) or {}
85
+ )
86
+ for k, v in d.items():
87
+ if k in ("all_of", "any_of"):
88
+ for c in v:
89
+ walk(c)
90
+ elif k == "not":
91
+ walk(v)
92
+ elif k == "within_ticks":
93
+ best.append(int(v))
94
+
95
+ walk(compiled.win_condition)
96
+ if best:
97
+ return max(1, min(best))
98
+ return max(1, compiled.scenario.termination.max_ticks)
99
+
100
+
101
  def _dimension_values(compiled: CompiledLevel, res: EpisodeResult) -> dict:
102
  """Map adapter signals -> Training reward dimensions, each in [0,1].
103
 
 
213
  def score_episode(compiled: CompiledLevel, res: EpisodeResult) -> ScoreCard:
214
  dims = _dimension_values(compiled, res)
215
  weights = _weights(compiled)
216
+ composite_base = _composite(dims, weights)
217
  pra = _pra_diagnostics(compiled, res, dims)
218
 
219
+ won = res.outcome == "win"
220
+ win_tick = int(res.signals.game_tick) if won else 0
221
+ win_turns = int(res.turns) if won else 0
222
+ budget = _win_budget(compiled)
223
+ speed = _clamp(1.0 - win_tick / budget) if won and win_tick > 0 else 0.0
224
+ composite = _clamp(composite_base + SPEED_BONUS * speed) if won else composite_base
225
+
226
  weakest = min(pra, key=pra.get)
227
  notes: list[str] = []
228
+ if won:
229
+ notes.append(
230
+ f"won in {win_turns} turns / tick {win_tick} of {budget} "
231
+ f"(speed {speed:.2f}, +{SPEED_BONUS * speed:.3f} bonus)"
232
+ )
233
  if res.outcome != "win":
234
  notes.append(f"objective not met ({res.outcome}); weakest link: {weakest}")
235
  if res.actions_issued and res.actions_warned / res.actions_issued > 0.25:
 
250
  weights=weights,
251
  weakest_link=weakest,
252
  notes=notes,
253
+ win_tick=win_tick,
254
+ win_turns=win_turns,
255
+ win_budget=budget,
256
+ speed=round(speed, 4),
257
+ composite_base=round(composite_base, 4),
258
  )
playback/run1/20260519-162229__qwen_qwen3.6-flash/_journal.jsonl ADDED
@@ -0,0 +1 @@
 
 
1
+ {"cell": "action-sequenced-execution:hard", "capability": "action", "split": "public", "seed": 1, "outcome": "loss", "composite": 0.1773, "perception": 0.6844, "reasoning": 0.6737, "action": 1.0, "weakest_link": "reasoning", "objective_progress": 0.375, "reward_vector": {"economy": 0.5, "military": 0.0, "territory": 0.5491, "scouting": 0.6, "objective": 0.375}, "turns": 41, "notes": ["objective not met (loss); weakest link: reasoning"], "_key": "action-sequenced-execution|hard|public|1"}
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/manifest.json ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scenario": "action-sequenced-execution:hard",
3
+ "pack_id": "action-sequenced-execution",
4
+ "level": "hard",
5
+ "capability": "action",
6
+ "run_id": "20260519-162229",
7
+ "model": "qwen/qwen3.6-flash",
8
+ "seed": 1,
9
+ "outcome": "loss",
10
+ "turns": 41,
11
+ "max_turns": 70,
12
+ "actions_issued": 78,
13
+ "actions_warned": 0,
14
+ "agent_stats": {
15
+ "turns": 41,
16
+ "tool_calls": 78,
17
+ "empty_replies": 0
18
+ },
19
+ "objective_progress": 0.375,
20
+ "reward_vector": {
21
+ "economy": 0.5,
22
+ "military": 0.0,
23
+ "territory": 0.5491,
24
+ "scouting": 0.6,
25
+ "objective": 0.375
26
+ },
27
+ "signals": {
28
+ "economy_value": 5000,
29
+ "explored_percent": 54.91,
30
+ "units_killed": 0,
31
+ "units_lost": 2
32
+ }
33
+ }
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/messages.json ADDED
The diff for this file is too large to render. See raw diff
 
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn001.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn002.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn003.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn004.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn005.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn006.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn007.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn008.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn009.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn010.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn011.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn012.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn013.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn014.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn015.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn016.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn017.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn018.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn019.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn020.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn021.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn022.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn023.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn024.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn025.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn026.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn027.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn028.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn029.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn030.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn031.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn032.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn033.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn034.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn035.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn036.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn037.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn038.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn039.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn040.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn041.png ADDED
playback/run1/20260519-162229__qwen_qwen3.6-flash/action-sequenced-execution:hard:public/seed1/minimap_turn042.png ADDED