python-gameplay-stuck-detection-screen-oscillation
Root cause (T67 run_v10): the stuck guard counted consecutive cycles where the screen type stayed the same. When the AI presses A at an NPC/sign, the game bounces dialog -> overworld -> dialog -> ... while the player never leaves the tile. Because the screen changes every cycle, _same_screen_count resets to 0 each time and the 40-cycle guard never fires — the controller loops forever.
Fix: key the consecutive-cycle counter on the player's tile (position unchanged), which is what "actually stuck" means, and ignore screen type. Screen oscillation is expected when interacting; tile stasis is not.
stuck_detector.py — the fixed detector (buggy one kept alongside for regression):
class SameTileStuckDetector:
"""FIXED: counts consecutive cycles where the player's *tile* is unchanged."""
def __init__(self, threshold: int = DEFAULT_STUCK_THRESHOLD): # 40
self.threshold = threshold
self._last_tile: Optional[tuple[int, int]] = None
self._same_tile_count = 0
self.fired = False
def update(self, state: PlayerState) -> bool:
tile = state.tile
if tile is None: # telemetry gap: skip, don't count
return self._is_stuck()
if tile == self._last_tile: # same TILE -> increment
self._same_tile_count += 1
else: # tile changed -> real progress
self._same_tile_count = 0
self._last_tile = tile
if self._is_stuck():
self.fired = True
return self._is_stuck()
def _is_stuck(self) -> bool:
return self._same_tile_count >= self.threshold
def reset(self) -> None: # call after recovery action
self._last_tile = None
self._same_tile_count = 0
self.fired = False
Controller integration in the E2E loop (T67 fixed path):
det = SameTileStuckDetector(threshold=40)
for state in controller.observe(): # tile=(x, y), screen=...
if det.update(state): # True even while screen flips
controller.press(B) # close dialog / cancel interaction
controller.move_to_random_adjacent_tile()
det.reset() # fresh start after recovery
Files: ~/stuck_detector.py, ~/test_stuck_detector.py (10 tests).
`python3 -m pytest test_stuck_detector.py -v` → **10 passed** (Python 3.14.4 / pytest 9.0.2). Direct T67 reproduction with the 40-cycle guard: ``` BUGGY (same-screen): fired_at = None # never fires across 45 oscillating cycles FIXED (same-tile): fired_at = 40 # trips at the 40th repeat walking 200 tiles, false positive = False # genuine progress never trips ``` Edge cases covered (all passing): 1. **T67 oscillation regression** — 45 dialog↔overworld cycles on one tile: buggy detector never fires (counter ≤ 1), fixed detector fires. 2. **Menu oscillation** — same-tile bounce with menu/overworld also trips the guard. 3. **No false positive while moving** — 200 cycles of real walking never trip. 4. **Counter resets on movement** — walking 10 tiles then oscillating fires only after the threshold from the new tile. 5. **Threshold boundary** — with threshold=40, the 39th repeat does not fire; the 40th repeat does (first observation establishes the tile). 6. **Small thresholds** — threshold=1 fires exactly on the first repeat. 7. **`None` tile telemetry** — 100 failed position reads produce zero stuck events. 8. **Telemetry gap mid-stall** — a `None` cycle pauses counting but does not wipe a real stall. 9. **Recovery loop** — after firing, moving to a fresh tile reports unstuck immediately; `reset()` clears the latch. ---
{"model": "deepseek-v4-flash", "problem_class": "python-gameplay-stuck-detection-screen-oscillation", "result": "passed", "tests": 10}The root cause: the stuck detector keyed its streak counter on the screen value (dialog/overworld/battle). A dialog → overworld → dialog → … loop at one tile flips the screen every frame, so the same-screen streak resets to 1 each cycle — the counter is structurally blind to any loop with ≥2 states. The fix replaces screen-equality with tile-equality and adds a starter-selection branch plus milestone events.
Implemented in ~/pokemon_stuck_fix/stuck_detector.py:
def _step(self, f: FrameState, cycle: int) -> Optional[StuckEvent]:
tile = f.tile # (map_group, map_id, x, y)
tile_changed = tile != self._last_tile
# 1. same-TILE streak — screen-agnostic, so dialog<->overworld loops
# still accumulate (the core fix)
if tile_changed:
recovered = self._stuck and self.cfg.emit_recovery
self._tile_streak = 1
self._no_progress = 0 # movement is progress
self._last_tile = tile
if recovered:
self._stuck = False
return StuckEvent("recovered", f"moved to tile {tile}", cycle, tile)
else:
self._tile_streak += 1
# 2. legacy same-SCREEN guard (also reset by movement) — keeps catching
# battle stalls
if tile_changed or f.screen != self._last_screen:
self._screen_streak = 1
else:
self._screen_streak += 1
self._last_screen = f.screen
# 3. milestone: party count ($D163) increased -> emit event, reset streaks
if self._last_party is None:
self._last_party = f.party_count
elif f.party_count > self._last_party:
ev = StuckEvent("milestone",
f"party_count {self._last_party}->{f.party_count} ($D163)",
cycle, tile)
self._milestones.append(ev)
self._reset_streaks(tile)
if f.party_count >= 1:
self._phase = Phase.EXPLORE
self._last_party = f.party_count
return ev # milestone clears any stuck latch
self._phase = Phase.STARTER if f.party_count == 0 else Phase.EXPLORE
return (self._starter_step(f, cycle, tile) if self._phase == Phase.STARTER
else self._explore_step(f, cycle, tile))
The starter-selection branch prevents false positives during the Oak intro (tile legitimately fixed for 100+ cycles while text scrolls): while $D163 == 0 it replaces the raw tile streak with a no-dialog-progress counter that resets whenever the dialog text pointer advances, still tripping at the same limit when nothing progresses:
def _starter_step(self, f, cycle, tile):
if (f.dialog_progress is not None and self._last_dialog is not None
and f.dialog_progress != self._last_dialog):
self._no_progress = 0
else:
self._no_progress += 1
self._last_dialog = f.dialog_progress
if self._no_progress >= self.cfg.starter_no_progress_limit + self.cfg.starter_grace:
return self._declare_stuck(cycle, tile,
f"starter selection: no dialog progress for {self._no_progress} cycles")
return None
The milestone event (party_count 0→1 = starter obtained, 1→2 = catch, …) emits an event usable for sparse-reward shaping, clears the stuck latch, and resets all streak counters so a fresh phase of play isn't punished for its first stationary seconds. WRAM reads are wrapped ($D163 party count, $D35D map group, $D35E map id, $D362/$D361 x/y), supporting both pyboy.memory[addr] and get_memory_value(bank, off).
Ran the built-in verification suite (`python3 stuck_detector.py`, 7 scenario groups, **15/15 PASS, exit 0**): - **A — the reported bug:** 46 cycles of `dialog↔overworld` on one fixed tile. Legacy screen-only detector: **blind** (`stuck=False` after 46 cycles). New detector: **stuck at cycle 30** via the same-TILE counter. ✅ - **B — same oscillation during starter selection** (`$D163==0`, no dialog progress): caught at ≤46 cycles by the no-progress branch. ✅ - **C — legitimate starter selection:** 60 cycles of fixed tile with advancing dialog text → no false stuck; at cycle 60 `$D163` 0→1 → `milestone` event, phase → EXPLORE, streaks reset; detection then correctly resumes (stuck @91 on a static tile after the milestone, proving the milestone doesn't permanently disable the guard). ✅ - **D — 200 cycles of normal exploration:** zero events. ✅ - **E — battle screen stall:** legacy same-screen guard still fires at cycle 120 (tile limit disabled). ✅ - **F — recovery:** moving after a stuck latch emits `recovered` and clears the latch. ✅ - **G — catch milestone:** `$D163` 1→2 at cycle 5 resets the tile streak; no stuck at 34, stuck exactly at 35 (limit counts from milestone). ✅ Edge cases exercised: idempotent stuck emission (one `stuck` per episode), tile-change recovery, screen flips mid-loop, pre-starter vs post-starter phases, bank-safe tile keys (map_group), unobservable dialog progress (falls back to no-progress counting), and both PyBoy memory-access styles via `GameReader` smoke test (OK).
{"model": "gpt-5.6-sol", "problem_class": "python-gameplay-stuck-detection-screen-oscillation", "result": "passed", "tests": 15}Root cause (T67 run_v10): the stuck guard counted consecutive cycles where the screen type stayed the same. When the AI presses A at an NPC/sign, the game bounces dialog -> overworld -> dialog -> ... while the player never leaves the tile. Because the screen changes every cycle, _same_screen_count resets to 0 each time and the 40-cycle guard never fires — the controller loops forever.
Fix: key the consecutive-cycle counter on the player's tile (position unchanged), which is what "actually stuck" means, and ignore screen type. Screen oscillation is expected when interacting; tile stasis is not.
stuck_detector.py — the fixed detector (buggy one kept alongside for regression):
class SameTileStuckDetector:
"""FIXED: counts consecutive cycles where the player's *tile* is unchanged."""
def __init__(self, threshold: int = DEFAULT_STUCK_THRESHOLD): # 40
self.threshold = threshold
self._last_tile: Optional[tuple[int, int]] = None
self._same_tile_count = 0
self.fired = False
def update(self, state: PlayerState) -> bool:
tile = state.tile
if tile is None: # telemetry gap: skip, don't count
return self._is_stuck()
if tile == self._last_tile: # same TILE -> increment
self._same_tile_count += 1
else: # tile changed -> real progress
self._same_tile_count = 0
self._last_tile = tile
if self._is_stuck():
self.fired = True
return self._is_stuck()
def _is_stuck(self) -> bool:
return self._same_tile_count >= self.threshold
def reset(self) -> None: # call after recovery action
self._last_tile = None
self._same_tile_count = 0
self.fired = False
Controller integration in the E2E loop (T67 fixed path):
det = SameTileStuckDetector(threshold=40)
for state in controller.observe(): # tile=(x, y), screen=...
if det.update(state): # True even while screen flips
controller.press(B) # close dialog / cancel interaction
controller.move_to_random_adjacent_tile()
det.reset() # fresh start after recovery
Files: ~/stuck_detector.py, ~/test_stuck_detector.py (10 tests).
`python3 -m pytest test_stuck_detector.py -v` → **10 passed** (Python 3.14.4 / pytest 9.0.2). Direct T67 reproduction with the 40-cycle guard: ``` BUGGY (same-screen): fired_at = None # never fires across 45 oscillating cycles FIXED (same-tile): fired_at = 40 # trips at the 40th repeat walking 200 tiles, false positive = False # genuine progress never trips ``` Edge cases covered (all passing): 1. **T67 oscillation regression** — 45 dialog↔overworld cycles on one tile: buggy detector never fires (counter ≤ 1), fixed detector fires. 2. **Menu oscillation** — same-tile bounce with menu/overworld also trips the guard. 3. **No false positive while moving** — 200 cycles of real walking never trip. 4. **Counter resets on movement** — walking 10 tiles then oscillating fires only after the threshold from the new tile. 5. **Threshold boundary** — with threshold=40, the 39th repeat does not fire; the 40th repeat does (first observation establishes the tile). 6. **Small thresholds** — threshold=1 fires exactly on the first repeat. 7. **`None` tile telemetry** — 100 failed position reads produce zero stuck events. 8. **Telemetry gap mid-stall** — a `None` cycle pauses counting but does not wipe a real stall. 9. **Recovery loop** — after firing, moving to a fresh tile reports unstuck immediately; `reset()` clears the latch. ---
{"model": "deepseek-v4-flash", "problem_class": "python-gameplay-stuck-detection-screen-oscillation", "result": "passed", "tests": 10}The root cause: the stuck detector keyed its streak counter on the screen value (dialog/overworld/battle). A dialog → overworld → dialog → … loop at one tile flips the screen every frame, so the same-screen streak resets to 1 each cycle — the counter is structurally blind to any loop with ≥2 states. The fix replaces screen-equality with tile-equality and adds a starter-selection branch plus milestone events.
Implemented in ~/pokemon_stuck_fix/stuck_detector.py:
def _step(self, f: FrameState, cycle: int) -> Optional[StuckEvent]:
tile = f.tile # (map_group, map_id, x, y)
tile_changed = tile != self._last_tile
# 1. same-TILE streak — screen-agnostic, so dialog<->overworld loops
# still accumulate (the core fix)
if tile_changed:
recovered = self._stuck and self.cfg.emit_recovery
self._tile_streak = 1
self._no_progress = 0 # movement is progress
self._last_tile = tile
if recovered:
self._stuck = False
return StuckEvent("recovered", f"moved to tile {tile}", cycle, tile)
else:
self._tile_streak += 1
# 2. legacy same-SCREEN guard (also reset by movement) — keeps catching
# battle stalls
if tile_changed or f.screen != self._last_screen:
self._screen_streak = 1
else:
self._screen_streak += 1
self._last_screen = f.screen
# 3. milestone: party count ($D163) increased -> emit event, reset streaks
if self._last_party is None:
self._last_party = f.party_count
elif f.party_count > self._last_party:
ev = StuckEvent("milestone",
f"party_count {self._last_party}->{f.party_count} ($D163)",
cycle, tile)
self._milestones.append(ev)
self._reset_streaks(tile)
if f.party_count >= 1:
self._phase = Phase.EXPLORE
self._last_party = f.party_count
return ev # milestone clears any stuck latch
self._phase = Phase.STARTER if f.party_count == 0 else Phase.EXPLORE
return (self._starter_step(f, cycle, tile) if self._phase == Phase.STARTER
else self._explore_step(f, cycle, tile))
The starter-selection branch prevents false positives during the Oak intro (tile legitimately fixed for 100+ cycles while text scrolls): while $D163 == 0 it replaces the raw tile streak with a no-dialog-progress counter that resets whenever the dialog text pointer advances, still tripping at the same limit when nothing progresses:
def _starter_step(self, f, cycle, tile):
if (f.dialog_progress is not None and self._last_dialog is not None
and f.dialog_progress != self._last_dialog):
self._no_progress = 0
else:
self._no_progress += 1
self._last_dialog = f.dialog_progress
if self._no_progress >= self.cfg.starter_no_progress_limit + self.cfg.starter_grace:
return self._declare_stuck(cycle, tile,
f"starter selection: no dialog progress for {self._no_progress} cycles")
return None
The milestone event (party_count 0→1 = starter obtained, 1→2 = catch, …) emits an event usable for sparse-reward shaping, clears the stuck latch, and resets all streak counters so a fresh phase of play isn't punished for its first stationary seconds. WRAM reads are wrapped ($D163 party count, $D35D map group, $D35E map id, $D362/$D361 x/y), supporting both pyboy.memory[addr] and get_memory_value(bank, off).
Ran the built-in verification suite (`python3 stuck_detector.py`, 7 scenario groups, **15/15 PASS, exit 0**): - **A — the reported bug:** 46 cycles of `dialog↔overworld` on one fixed tile. Legacy screen-only detector: **blind** (`stuck=False` after 46 cycles). New detector: **stuck at cycle 30** via the same-TILE counter. ✅ - **B — same oscillation during starter selection** (`$D163==0`, no dialog progress): caught at ≤46 cycles by the no-progress branch. ✅ - **C — legitimate starter selection:** 60 cycles of fixed tile with advancing dialog text → no false stuck; at cycle 60 `$D163` 0→1 → `milestone` event, phase → EXPLORE, streaks reset; detection then correctly resumes (stuck @91 on a static tile after the milestone, proving the milestone doesn't permanently disable the guard). ✅ - **D — 200 cycles of normal exploration:** zero events. ✅ - **E — battle screen stall:** legacy same-screen guard still fires at cycle 120 (tile limit disabled). ✅ - **F — recovery:** moving after a stuck latch emits `recovered` and clears the latch. ✅ - **G — catch milestone:** `$D163` 1→2 at cycle 5 resets the tile streak; no stuck at 34, stuck exactly at 35 (limit counts from milestone). ✅ Edge cases exercised: idempotent stuck emission (one `stuck` per episode), tile-change recovery, screen flips mid-loop, pre-starter vs post-starter phases, bank-safe tile keys (map_group), unobservable dialog progress (falls back to no-progress counting), and both PyBoy memory-access styles via `GameReader` smoke test (OK).
{"model": "gpt-5.6-sol", "problem_class": "python-gameplay-stuck-detection-screen-oscillation", "result": "passed", "tests": 15}