Problem class: python-gameplay-de-hardcoding-jev-routing
I built and verified a self-contained reference implementation at ~/jev_routing_solution/ and wrote the full solution to ~/SOLUTION.md. Verification: 13/13 tests pass.
Here is the solution document:
Problem class: python-gameplay-de-hardcoding-jev-routing
The agent took two gameplay decisions by fiat:
CHARMANDER (a fixed-tile walk to one ball), andselect_move(1).Those constants bypassed RAM, so no test could distinguish a working agent from a broken one: the test asserted the constant equalled itself. The fix is a typed choice-question layer that (a) builds a closed question from RAM state, (b) reads the ground truth from RAM, (c) lets the decision distribution pick a token, (d) translates the token into a tool call, and (e) logs the raw distribution and chosen action on every decision row. Recovery ladders stay in the transport and retry the same semantic call; they never choose.
2.1 The decision was a constant, not a function of state
# BEFORE (anti-pattern)
STARTER = "CHARMANDER"
def take_starter(self):
self.walk_to_fixed_tile() # always the same ball
self.press_a()
def battle(self):
self.select_move(1) # always slot 1
Nothing consulted the intro dialog (which names the species) or battle RAM (move slots, party, bag, run flag). The inputs existed but were never read.
2.2 The tests were tautological
# BEFORE (invalid test): asserts the constant again
def test_take_starter():
agent.take_starter()
assert agent.species == "CHARMANDER"
This passes when routing is broken, when RAM is wrong, and when the dialog names a different species. It is a change-detector, not a verification.
2.3 Decisions were not logged, so "what the model wanted" versus "what the code did" was unobservable.
2.4 Recovery mixed transport with decision-making: retry helpers re-picked highlighted entries on failure, silently changing the semantic action.
2.5 Missing abstractions: no typed ChoiceQuestion, no domain vocabulary, no token→tool translator, no RAM adapter. Adding these four seams is the entire fix.
jev_routing/
__init__.py
questions.py # typed vocabulary: ChoiceQuestion / options / enums
ram.py # RAM adapter + ground-truth parser
domain.py # RAM -> question builders (species, battle)
transport.py # ToolCall, Transport, RecoveryLadder (transport only)
router.py # distribution -> token -> tool call -> DecisionRow log
hardcoded.py # the retired path, used only by the validity guard
tests/
fakes.py
test_routing.py
questions.pyfrom __future__ import annotations
from dataclasses import dataclass, field
from enum import Enum
from typing import Iterator, Mapping
class QuestionKind(str, Enum):
SPECIES_CHOICE = "species_choice"
BATTLE_ACTION = "battle_action"
class Species(str, Enum):
BULBASAUR = "BULBASAUR"
CHARMANDER = "CHARMANDER"
SQUIRTLE = "SQUIRTLE"
class BattleAction(str, Enum):
MOVE_1 = "MOVE_1"; MOVE_2 = "MOVE_2"; MOVE_3 = "MOVE_3"; MOVE_4 = "MOVE_4"
SWITCH = "SWITCH"; ITEM = "ITEM"; RUN = "RUN"
@dataclass(frozen=True)
class ChoiceOption:
token: str
label: str
enabled: bool = True
@dataclass(frozen=True)
class ChoiceQuestion:
kind: QuestionKind
prompt: str
options: tuple[ChoiceOption, ...]
ground_truth: str | None = None
context: Mapping[str, str] = field(default_factory=dict)
def __post_init__(self) -> None:
tokens = [o.token for o in self.options]
if len(tokens) != len(set(tokens)):
raise ValueError(f"duplicate option tokens in {self.kind}: {tokens}")
if not tokens:
raise ValueError(f"{self.kind} question has no options")
def tokens(self) -> tuple[str, ...]:
return tuple(o.token for o in self.options if o.enabled)
def __iter__(self) -> Iterator[ChoiceOption]:
return iter(self.options)
def by_token(self, token: str) -> ChoiceOption:
for option in self.options:
if option.token == token:
return option
raise KeyError(f"{token!r} is not a choice for {self.kind}")
ram.pyfrom __future__ import annotations
SPECIES_NAMES: tuple[str, ...] = ("BULBASAUR", "CHARMANDER", "SQUIRTLE")
DIALOG_PTR_ADDR = 0xD000
DIALOG_BUF_ADDR = 0xD100
DIALOG_TERMINATOR = 0x00
MOVE_NAMES_ADDR = 0xD200
MOVE_SLOT_LEN = 13
PARTY_COUNT_ADDR = 0xD240
BAG_ITEM_COUNT_ADDR = 0xD241
BATTLE_FLAGS_ADDR = 0xD242 # bit0 = can run, bit1 = can switch
class GameRAM:
def __init__(self, size: int = 0x10000) -> None:
self._mem = bytearray(size)
def write(self, addr: int, data: bytes | bytearray) -> None:
self._mem[addr:addr + len(data)] = data
def read(self, addr: int, length: int) -> bytes:
return bytes(self._mem[addr:addr + length])
def read_u8(self, addr: int) -> int:
return self._mem[addr]
def read_u16(self, addr: int) -> int:
return self._mem[addr] | (self._mem[addr + 1] << 8)
def write_u16(self, addr: int, value: int) -> None:
self._mem[addr] = value & 0xFF
self._mem[addr + 1] = (value >> 8) & 0xFF
def write_c_string(self, addr: int, text: str, width: int | None = None) -> None:
raw = text.encode("ascii", "replace") + b"\x00"
if width is not None:
raw = raw[:width].ljust(width, b"\x00")
self._mem[addr:addr + len(raw)] = raw
def read_c_string(self, addr: int, terminator: int = DIALOG_TERMINATOR) -> str:
out = bytearray()
cursor = addr
while self._mem[cursor] != terminator:
out.append(self._mem[cursor])
cursor += 1
if cursor - addr > 4096:
raise ValueError("unterminated RAM string")
return out.decode("ascii", "replace")
def read_dialog_text(self) -> str:
return self.read_c_string(self.read_u16(DIALOG_PTR_ADDR))
def read_move_names(self) -> list[str]:
return [self.read_c_string(MOVE_NAMES_ADDR + s * MOVE_SLOT_LEN)
for s in range(4)]
def read_party_count(self) -> int:
return self.read_u8(PARTY_COUNT_ADDR)
def read_item_count(self) -> int:
return self.read_u8(BAG_ITEM_COUNT_ADDR)
def read_battle_flags(self) -> int:
return self.read_u8(BATTLE_FLAGS_ADDR)
def parse_species_from_dialog(dialog: str) -> str | None:
upper = dialog.upper()
for name in SPECIES_NAMES:
if name in upper:
return name
return None
domain.pyfrom __future__ import annotations
from .questions import BattleAction, ChoiceOption, ChoiceQuestion, QuestionKind
from .ram import GameRAM, MOVE_NAMES_ADDR, MOVE_SLOT_LEN, parse_species_from_dialog
STARTER_BALLS: tuple[str, ...] = ("BALL_LEFT", "BALL_CENTER", "BALL_RIGHT")
_BALL_TO_SPECIES = dict(zip(STARTER_BALLS, ("BULBASAUR", "CHARMANDER", "SQUIRTLE")))
def build_species_question(ram: GameRAM) -> ChoiceQuestion:
dialog = ram.read_dialog_text()
truth = parse_species_from_dialog(dialog)
options = tuple(
ChoiceOption(token=species, label=f"{ball} -> {species}")
for ball, species in _BALL_TO_SPECIES.items()
)
return ChoiceQuestion(
kind=QuestionKind.SPECIES_CHOICE,
prompt=f"Which starter will you take? (RAM dialog: {dialog!r})",
options=options,
ground_truth=truth,
context={"dialog": dialog, "balls": ",".join(STARTER_BALLS)},
)
_MOVE_TOKENS = (BattleAction.MOVE_1, BattleAction.MOVE_2, BattleAction.MOVE_3, BattleAction.MOVE_4)
_AUX_TOKENS = (BattleAction.SWITCH, BattleAction.ITEM, BattleAction.RUN)
def build_battle_question(ram: GameRAM) -> ChoiceQuestion:
move_names = ram.read_move_names()
party = ram.read_party_count()
items = ram.read_item_count()
flags = ram.read_battle_flags()
can_switch, can_run, can_item = party > 1, bool(flags & 0x01), items > 0
options: list[ChoiceOption] = []
for idx, (token, name) in enumerate(zip(_MOVE_TOKENS, move_names), start=1):
usable = bool(name.strip()) and not name.startswith("\x00")
options.append(ChoiceOption(token=token.value,
label=f"MOVE_{idx}: {name.strip() or '---'}",
enabled=usable))
options.append(ChoiceOption(token=BattleAction.SWITCH.value, label="SWITCH", enabled=can_switch))
options.append(ChoiceOption(token=BattleAction.ITEM.value, label="ITEM", enabled=can_item))
options.append(ChoiceOption(token=BattleAction.RUN.value, label="RUN", enabled=can_run))
return ChoiceQuestion(
kind=QuestionKind.BATTLE_ACTION,
prompt="What should the active Pokemon do?",
options=tuple(options),
ground_truth=None,
context={"moves": "|".join(move_names), "party": str(party),
"items": str(items), "flags": str(flags)},
)
transport.pyfrom __future__ import annotations
from dataclasses import dataclass, field
from typing import Callable, Protocol
@dataclass(frozen=True)
class ToolCall:
name: str
arguments: dict = field(default_factory=dict)
def key(self) -> tuple[str, tuple[tuple[str, object], ...]]:
return (self.name, tuple(sorted(self.arguments.items())))
@dataclass
class TransportResult:
ok: bool
calls: list[ToolCall]
attempts: int
recovered: bool = False
class RecoveryLadder:
"""Ordered neutral inputs used only to unwedge the transport."""
def __init__(self, steps: list[Callable[[], None]] | None = None) -> None:
self.steps = list(steps or [])
self.invocations = 0
def recover(self) -> None:
self.invocations += 1
for step in self.steps:
step()
class Environment(Protocol):
def execute(self, call: ToolCall) -> bool: ...
class Transport:
"""Executes one semantic ToolCall with a transport-only retry loop."""
def __init__(self, environment: Environment, recovery: RecoveryLadder | None = None,
max_attempts: int = 3) -> None:
self.environment = environment
self.recovery = recovery or RecoveryLadder()
self.max_attempts = max_attempts
self.history: list[ToolCall] = []
def execute(self, call: ToolCall) -> TransportResult:
attempts, recovered = 0, False
while attempts < self.max_attempts:
attempts += 1
self.history.append(call) # same call every attempt
if self.environment.execute(call):
return TransportResult(True, [call], attempts, recovered)
recovered = True
self.recovery.recover()
return TransportResult(False, [call], attempts, recovered)
router.pyfrom __future__ import annotations
import random
from dataclasses import dataclass, field
from typing import Callable, Mapping
from .questions import QuestionKind
from .transport import ToolCall, Transport
Translator = Callable[["DecisionRouter", object, str], ToolCall]
@dataclass
class DecisionRow:
decision_id: int
kind: str
prompt: str
ground_truth: str | None
distribution: dict[str, float]
chosen: str
tool_call: ToolCall
transport_ok: bool
attempts: int
recovered: bool
def as_dict(self) -> dict:
return {
"decision_id": self.decision_id, "kind": self.kind, "prompt": self.prompt,
"ground_truth": self.ground_truth, "distribution": dict(self.distribution),
"chosen": self.chosen,
"tool_call": {"name": self.tool_call.name,
"arguments": dict(self.tool_call.arguments)},
"transport_ok": self.transport_ok, "attempts": self.attempts,
"recovered": self.recovered,
}
def _translate_species(router, question, token):
return ToolCall("choose_starter", {"species": token})
def _translate_battle(router, question, token):
if token.startswith("MOVE_"):
return ToolCall("select_move", {"slot": int(token.split("_")[1])})
return {"SWITCH": ToolCall("open_party", {}), "ITEM": ToolCall("open_bag", {}),
"RUN": ToolCall("run", {})}[token]
DEFAULT_TRANSLATORS = {QuestionKind.SPECIES_CHOICE: _translate_species,
QuestionKind.BATTLE_ACTION: _translate_battle}
class DecisionRouter:
def __init__(self, transport, translators=None, policy="argmax", rng=None):
self.transport = transport
self.translators = dict(DEFAULT_TRANSLATORS if translators is None else translators)
self.policy = policy
self.rng = rng or random.Random(0)
self.log: list[DecisionRow] = []
def _pick(self, tokens, distribution):
if self.policy == "argmax":
return tokens[min(range(len(tokens)),
key=lambda i: (-distribution[tokens[i]], i))]
if self.policy == "sample":
weights = [max(0.0, distribution[t]) for t in tokens]
if sum(weights) <= 0:
raise ValueError("distribution has zero mass")
return self.rng.choices(list(tokens), weights=weights, k=1)[0]
raise ValueError(f"unknown policy {self.policy!r}")
def decide(self, question, distribution) -> DecisionRow:
tokens = question.tokens()
missing = [t for t in tokens if t not in distribution]
extra = [t for t in distribution if t not in tokens]
if missing or extra:
raise ValueError(f"distribution must be defined over enabled tokens; "
f"missing={missing} extra={extra}")
chosen = self._pick(tokens, distribution)
call = self.translators[question.kind](self, question, chosen)
result = self.transport.execute(call)
row = DecisionRow(len(self.log), question.kind.value, question.prompt,
question.ground_truth,
{t: float(distribution[t]) for t in tokens}, chosen, call,
result.ok, result.attempts, result.recovered)
self.log.append(row)
return row
# starter
q = build_species_question(ram)
dist = model.choose_distribution(q) # {"BULBASAUR": .1, "CHARMANDER": .1, "SQUIRTLE": .8}
row = router.decide(q, dist) # -> choose_starter(species=...)
# battle
q = build_battle_question(ram)
dist = model.choose_distribution(q) # over MOVE_1..4/SWITCH/ITEM/RUN
row = router.decide(q, dist) # -> select_move/open_party/open_bag/run
The walk to a ball is now transport for the chosen species token; it is no longer what selects the species.
cd ~/jev_routing_solution
python3 -m pytest tests/ -v
Result (Python 3.14.4, pytest 9.0.2): 13 passed
test_species_ground_truth_read_from_ram PASSED
test_battle_options_come_from_ram_state PASSED
test_species_distribution_selects_squirtle_not_hardcoded_charmander PASSED
test_species_alternate_distribution_flips_outcome PASSED
test_battle_distribution_selects_run_not_hardcoded_move_1 PASSED
test_battle_move_slot_translates_to_correct_tool_arguments PASSED
test_switch_and_item_use_their_own_tools PASSED
test_distribution_must_cover_enabled_tokens PASSED
test_disabled_tokens_are_not_routable PASSED
test_log_records_raw_distribution_and_action_per_row PASSED
test_recovery_ladder_retries_same_call_without_changing_decision PASSED
test_transport_exhaustion_reports_not_ok_without_substituting_action PASSED
test_hardcoded_path_is_rejected_by_the_forced_choice_test PASSED
SQUIRTLE ⇒ choose_starter(species="SQUIRTLE"), not the old CHARMANDER.RUN ⇒ ToolCall("run"), not select_move(1).test_species_alternate_distribution_flips_outcome re-asks the same question with a different distribution and asserts the tool call changes.test_hardcoded_path_is_rejected_by_the_forced_choice_test asserts the retired HardcodedAgent cannot produce SQUIRTLE or run, proving the forced-choice tests fail on the hardcoded path.test_species_ground_truth_read_from_ram writes "It's SQUIRTLE!" into the dialog buffer and points the pointer at it, asserting ground_truth == "SQUIRTLE". test_battle_options_come_from_ram_state clears a move slot and zeroes party/items/flags and asserts those tokens become non-routable.
test_recovery_ladder_retries_same_call_without_changing_decision forces one failure and asserts chosen == "RUN", the call is still run, the neutral step ran once, and transport.history == ["run", "run"]. Exhaustion surfaces transport_ok=False without substituting an action.
MOVE_1..4 / SWITCH / ITEM / RUN.Artifacts: ~/jev_routing_solution/ (runnable code + tests) and ~/SOLUTION.md (full document incl. appendices with __init__.py, hardcoded.py, fakes.py, and the complete test file). Verified with python3 -m pytest tests/ -v → 13 passed.
# Evidence - Problem class: python-gameplay-de-hardcoding-jev-routing - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T21:58:50.834Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Retiring hardcoded gameplay assists (fixed-tile starter walk, battle always select_move(1)) by routing decisions through a typed choice-question decision layer: add a domain-specific question set (species choice over the three balls; battle action over MOVE_1..4/SWITCH/ITEM/RUN vocabulary), read the ground truth from RAM (dialog text names the species), let the decision distribution pick the action, translate choice tokens to tool calls, and log raw distribution + action per decision row. Tests must force a different choice and assert a DIFFERENT outcome \u2014 a test that passes on the hardcoded path is invalid. Keep recovery ladders as transport, not decision-makers.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "python-gameplay-de-hardcoding-jev-routing", "provider": "openrouter", "solved_at": "2026-09-24T21:58:50.836Z", "version": ""}I built and verified a self-contained reference implementation at ~/jev_routing_solution/ and wrote the full solution to ~/SOLUTION.md. Verification: 13/13 tests pass.
Here is the solution document:
Problem class: python-gameplay-de-hardcoding-jev-routing
The agent took two gameplay decisions by fiat:
CHARMANDER (a fixed-tile walk to one ball), andselect_move(1).Those constants bypassed RAM, so no test could distinguish a working agent from a broken one: the test asserted the constant equalled itself. The fix is a typed choice-question layer that (a) builds a closed question from RAM state, (b) reads the ground truth from RAM, (c) lets the decision distribution pick a token, (d) translates the token into a tool call, and (e) logs the raw distribution and chosen action on every decision row. Recovery ladders stay in the transport and retry the same semantic call; they never choose.
2.1 The decision was a constant, not a function of state
# BEFORE (anti-pattern)
STARTER = "CHARMANDER"
def take_starter(self):
self.walk_to_fixed_tile() # always the same ball
self.press_a()
def battle(self):
self.select_move(1) # always slot 1
Nothing consulted the intro dialog (which names the species) or battle RAM (move slots, party, bag, run flag). The inputs existed but were never read.
2.2 The tests were tautological
# BEFORE (invalid test): asserts the constant again
def test_take_starter():
agent.take_starter()
assert agent.species == "CHARMANDER"
This passes when routing is broken, when RAM is wrong, and when the dialog names a different species. It is a change-detector, not a verification.
2.3 Decisions were not logged, so "what the model wanted" versus "what the code did" was unobservable.
2.4 Recovery mixed transport with decision-making: retry helpers re-picked highlighted entries on failure, silently changing the semantic action.
2.5 Missing abstractions: no typed ChoiceQuestion, no domain vocabulary, no token→tool translator, no RAM adapter. Adding these four seams is the entire fix.
jev_routing/
__init__.py
questions.py # typed vocabulary: ChoiceQuestion / options / enums
ram.py # RAM adapter + ground-truth parser
domain.py # RAM -> question builders (species, battle)
transport.py # ToolCall, Transport, RecoveryLadder (transport only)
router.py # distribution -> token -> tool call -> DecisionRow log
hardcoded.py # the retired path, used only by the validity guard
tests/
fakes.py
test_routing.py
questions.pyfrom __future__ import annotations
from dataclasses import dataclass, field
from enum import Enum
from typing import Iterator, Mapping
class QuestionKind(str, Enum):
SPECIES_CHOICE = "species_choice"
BATTLE_ACTION = "battle_action"
class Species(str, Enum):
BULBASAUR = "BULBASAUR"
CHARMANDER = "CHARMANDER"
SQUIRTLE = "SQUIRTLE"
class BattleAction(str, Enum):
MOVE_1 = "MOVE_1"; MOVE_2 = "MOVE_2"; MOVE_3 = "MOVE_3"; MOVE_4 = "MOVE_4"
SWITCH = "SWITCH"; ITEM = "ITEM"; RUN = "RUN"
@dataclass(frozen=True)
class ChoiceOption:
token: str
label: str
enabled: bool = True
@dataclass(frozen=True)
class ChoiceQuestion:
kind: QuestionKind
prompt: str
options: tuple[ChoiceOption, ...]
ground_truth: str | None = None
context: Mapping[str, str] = field(default_factory=dict)
def __post_init__(self) -> None:
tokens = [o.token for o in self.options]
if len(tokens) != len(set(tokens)):
raise ValueError(f"duplicate option tokens in {self.kind}: {tokens}")
if not tokens:
raise ValueError(f"{self.kind} question has no options")
def tokens(self) -> tuple[str, ...]:
return tuple(o.token for o in self.options if o.enabled)
def __iter__(self) -> Iterator[ChoiceOption]:
return iter(self.options)
def by_token(self, token: str) -> ChoiceOption:
for option in self.options:
if option.token == token:
return option
raise KeyError(f"{token!r} is not a choice for {self.kind}")
ram.pyfrom __future__ import annotations
SPECIES_NAMES: tuple[str, ...] = ("BULBASAUR", "CHARMANDER", "SQUIRTLE")
DIALOG_PTR_ADDR = 0xD000
DIALOG_BUF_ADDR = 0xD100
DIALOG_TERMINATOR = 0x00
MOVE_NAMES_ADDR = 0xD200
MOVE_SLOT_LEN = 13
PARTY_COUNT_ADDR = 0xD240
BAG_ITEM_COUNT_ADDR = 0xD241
BATTLE_FLAGS_ADDR = 0xD242 # bit0 = can run, bit1 = can switch
class GameRAM:
def __init__(self, size: int = 0x10000) -> None:
self._mem = bytearray(size)
def write(self, addr: int, data: bytes | bytearray) -> None:
self._mem[addr:addr + len(data)] = data
def read(self, addr: int, length: int) -> bytes:
return bytes(self._mem[addr:addr + length])
def read_u8(self, addr: int) -> int:
return self._mem[addr]
def read_u16(self, addr: int) -> int:
return self._mem[addr] | (self._mem[addr + 1] << 8)
def write_u16(self, addr: int, value: int) -> None:
self._mem[addr] = value & 0xFF
self._mem[addr + 1] = (value >> 8) & 0xFF
def write_c_string(self, addr: int, text: str, width: int | None = None) -> None:
raw = text.encode("ascii", "replace") + b"\x00"
if width is not None:
raw = raw[:width].ljust(width, b"\x00")
self._mem[addr:addr + len(raw)] = raw
def read_c_string(self, addr: int, terminator: int = DIALOG_TERMINATOR) -> str:
out = bytearray()
cursor = addr
while self._mem[cursor] != terminator:
out.append(self._mem[cursor])
cursor += 1
if cursor - addr > 4096:
raise ValueError("unterminated RAM string")
return out.decode("ascii", "replace")
def read_dialog_text(self) -> str:
return self.read_c_string(self.read_u16(DIALOG_PTR_ADDR))
def read_move_names(self) -> list[str]:
return [self.read_c_string(MOVE_NAMES_ADDR + s * MOVE_SLOT_LEN)
for s in range(4)]
def read_party_count(self) -> int:
return self.read_u8(PARTY_COUNT_ADDR)
def read_item_count(self) -> int:
return self.read_u8(BAG_ITEM_COUNT_ADDR)
def read_battle_flags(self) -> int:
return self.read_u8(BATTLE_FLAGS_ADDR)
def parse_species_from_dialog(dialog: str) -> str | None:
upper = dialog.upper()
for name in SPECIES_NAMES:
if name in upper:
return name
return None
domain.pyfrom __future__ import annotations
from .questions import BattleAction, ChoiceOption, ChoiceQuestion, QuestionKind
from .ram import GameRAM, MOVE_NAMES_ADDR, MOVE_SLOT_LEN, parse_species_from_dialog
STARTER_BALLS: tuple[str, ...] = ("BALL_LEFT", "BALL_CENTER", "BALL_RIGHT")
_BALL_TO_SPECIES = dict(zip(STARTER_BALLS, ("BULBASAUR", "CHARMANDER", "SQUIRTLE")))
def build_species_question(ram: GameRAM) -> ChoiceQuestion:
dialog = ram.read_dialog_text()
truth = parse_species_from_dialog(dialog)
options = tuple(
ChoiceOption(token=species, label=f"{ball} -> {species}")
for ball, species in _BALL_TO_SPECIES.items()
)
return ChoiceQuestion(
kind=QuestionKind.SPECIES_CHOICE,
prompt=f"Which starter will you take? (RAM dialog: {dialog!r})",
options=options,
ground_truth=truth,
context={"dialog": dialog, "balls": ",".join(STARTER_BALLS)},
)
_MOVE_TOKENS = (BattleAction.MOVE_1, BattleAction.MOVE_2, BattleAction.MOVE_3, BattleAction.MOVE_4)
_AUX_TOKENS = (BattleAction.SWITCH, BattleAction.ITEM, BattleAction.RUN)
def build_battle_question(ram: GameRAM) -> ChoiceQuestion:
move_names = ram.read_move_names()
party = ram.read_party_count()
items = ram.read_item_count()
flags = ram.read_battle_flags()
can_switch, can_run, can_item = party > 1, bool(flags & 0x01), items > 0
options: list[ChoiceOption] = []
for idx, (token, name) in enumerate(zip(_MOVE_TOKENS, move_names), start=1):
usable = bool(name.strip()) and not name.startswith("\x00")
options.append(ChoiceOption(token=token.value,
label=f"MOVE_{idx}: {name.strip() or '---'}",
enabled=usable))
options.append(ChoiceOption(token=BattleAction.SWITCH.value, label="SWITCH", enabled=can_switch))
options.append(ChoiceOption(token=BattleAction.ITEM.value, label="ITEM", enabled=can_item))
options.append(ChoiceOption(token=BattleAction.RUN.value, label="RUN", enabled=can_run))
return ChoiceQuestion(
kind=QuestionKind.BATTLE_ACTION,
prompt="What should the active Pokemon do?",
options=tuple(options),
ground_truth=None,
context={"moves": "|".join(move_names), "party": str(party),
"items": str(items), "flags": str(flags)},
)
transport.pyfrom __future__ import annotations
from dataclasses import dataclass, field
from typing import Callable, Protocol
@dataclass(frozen=True)
class ToolCall:
name: str
arguments: dict = field(default_factory=dict)
def key(self) -> tuple[str, tuple[tuple[str, object], ...]]:
return (self.name, tuple(sorted(self.arguments.items())))
@dataclass
class TransportResult:
ok: bool
calls: list[ToolCall]
attempts: int
recovered: bool = False
class RecoveryLadder:
"""Ordered neutral inputs used only to unwedge the transport."""
def __init__(self, steps: list[Callable[[], None]] | None = None) -> None:
self.steps = list(steps or [])
self.invocations = 0
def recover(self) -> None:
self.invocations += 1
for step in self.steps:
step()
class Environment(Protocol):
def execute(self, call: ToolCall) -> bool: ...
class Transport:
"""Executes one semantic ToolCall with a transport-only retry loop."""
def __init__(self, environment: Environment, recovery: RecoveryLadder | None = None,
max_attempts: int = 3) -> None:
self.environment = environment
self.recovery = recovery or RecoveryLadder()
self.max_attempts = max_attempts
self.history: list[ToolCall] = []
def execute(self, call: ToolCall) -> TransportResult:
attempts, recovered = 0, False
while attempts < self.max_attempts:
attempts += 1
self.history.append(call) # same call every attempt
if self.environment.execute(call):
return TransportResult(True, [call], attempts, recovered)
recovered = True
self.recovery.recover()
return TransportResult(False, [call], attempts, recovered)
router.pyfrom __future__ import annotations
import random
from dataclasses import dataclass, field
from typing import Callable, Mapping
from .questions import QuestionKind
from .transport import ToolCall, Transport
Translator = Callable[["DecisionRouter", object, str], ToolCall]
@dataclass
class DecisionRow:
decision_id: int
kind: str
prompt: str
ground_truth: str | None
distribution: dict[str, float]
chosen: str
tool_call: ToolCall
transport_ok: bool
attempts: int
recovered: bool
def as_dict(self) -> dict:
return {
"decision_id": self.decision_id, "kind": self.kind, "prompt": self.prompt,
"ground_truth": self.ground_truth, "distribution": dict(self.distribution),
"chosen": self.chosen,
"tool_call": {"name": self.tool_call.name,
"arguments": dict(self.tool_call.arguments)},
"transport_ok": self.transport_ok, "attempts": self.attempts,
"recovered": self.recovered,
}
def _translate_species(router, question, token):
return ToolCall("choose_starter", {"species": token})
def _translate_battle(router, question, token):
if token.startswith("MOVE_"):
return ToolCall("select_move", {"slot": int(token.split("_")[1])})
return {"SWITCH": ToolCall("open_party", {}), "ITEM": ToolCall("open_bag", {}),
"RUN": ToolCall("run", {})}[token]
DEFAULT_TRANSLATORS = {QuestionKind.SPECIES_CHOICE: _translate_species,
QuestionKind.BATTLE_ACTION: _translate_battle}
class DecisionRouter:
def __init__(self, transport, translators=None, policy="argmax", rng=None):
self.transport = transport
self.translators = dict(DEFAULT_TRANSLATORS if translators is None else translators)
self.policy = policy
self.rng = rng or random.Random(0)
self.log: list[DecisionRow] = []
def _pick(self, tokens, distribution):
if self.policy == "argmax":
return tokens[min(range(len(tokens)),
key=lambda i: (-distribution[tokens[i]], i))]
if self.policy == "sample":
weights = [max(0.0, distribution[t]) for t in tokens]
if sum(weights) <= 0:
raise ValueError("distribution has zero mass")
return self.rng.choices(list(tokens), weights=weights, k=1)[0]
raise ValueError(f"unknown policy {self.policy!r}")
def decide(self, question, distribution) -> DecisionRow:
tokens = question.tokens()
missing = [t for t in tokens if t not in distribution]
extra = [t for t in distribution if t not in tokens]
if missing or extra:
raise ValueError(f"distribution must be defined over enabled tokens; "
f"missing={missing} extra={extra}")
chosen = self._pick(tokens, distribution)
call = self.translators[question.kind](self, question, chosen)
result = self.transport.execute(call)
row = DecisionRow(len(self.log), question.kind.value, question.prompt,
question.ground_truth,
{t: float(distribution[t]) for t in tokens}, chosen, call,
result.ok, result.attempts, result.recovered)
self.log.append(row)
return row
# starter
q = build_species_question(ram)
dist = model.choose_distribution(q) # {"BULBASAUR": .1, "CHARMANDER": .1, "SQUIRTLE": .8}
row = router.decide(q, dist) # -> choose_starter(species=...)
# battle
q = build_battle_question(ram)
dist = model.choose_distribution(q) # over MOVE_1..4/SWITCH/ITEM/RUN
row = router.decide(q, dist) # -> select_move/open_party/open_bag/run
The walk to a ball is now transport for the chosen species token; it is no longer what selects the species.
cd ~/jev_routing_solution
python3 -m pytest tests/ -v
Result (Python 3.14.4, pytest 9.0.2): 13 passed
test_species_ground_truth_read_from_ram PASSED
test_battle_options_come_from_ram_state PASSED
test_species_distribution_selects_squirtle_not_hardcoded_charmander PASSED
test_species_alternate_distribution_flips_outcome PASSED
test_battle_distribution_selects_run_not_hardcoded_move_1 PASSED
test_battle_move_slot_translates_to_correct_tool_arguments PASSED
test_switch_and_item_use_their_own_tools PASSED
test_distribution_must_cover_enabled_tokens PASSED
test_disabled_tokens_are_not_routable PASSED
test_log_records_raw_distribution_and_action_per_row PASSED
test_recovery_ladder_retries_same_call_without_changing_decision PASSED
test_transport_exhaustion_reports_not_ok_without_substituting_action PASSED
test_hardcoded_path_is_rejected_by_the_forced_choice_test PASSED
SQUIRTLE ⇒ choose_starter(species="SQUIRTLE"), not the old CHARMANDER.RUN ⇒ ToolCall("run"), not select_move(1).test_species_alternate_distribution_flips_outcome re-asks the same question with a different distribution and asserts the tool call changes.test_hardcoded_path_is_rejected_by_the_forced_choice_test asserts the retired HardcodedAgent cannot produce SQUIRTLE or run, proving the forced-choice tests fail on the hardcoded path.test_species_ground_truth_read_from_ram writes "It's SQUIRTLE!" into the dialog buffer and points the pointer at it, asserting ground_truth == "SQUIRTLE". test_battle_options_come_from_ram_state clears a move slot and zeroes party/items/flags and asserts those tokens become non-routable.
test_recovery_ladder_retries_same_call_without_changing_decision forces one failure and asserts chosen == "RUN", the call is still run, the neutral step ran once, and transport.history == ["run", "run"]. Exhaustion surfaces transport_ok=False without substituting an action.
MOVE_1..4 / SWITCH / ITEM / RUN.Artifacts: ~/jev_routing_solution/ (runnable code + tests) and ~/SOLUTION.md (full document incl. appendices with __init__.py, hardcoded.py, fakes.py, and the complete test file). Verified with python3 -m pytest tests/ -v → 13 passed.
# Evidence - Problem class: python-gameplay-de-hardcoding-jev-routing - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T21:58:50.834Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Retiring hardcoded gameplay assists (fixed-tile starter walk, battle always select_move(1)) by routing decisions through a typed choice-question decision layer: add a domain-specific question set (species choice over the three balls; battle action over MOVE_1..4/SWITCH/ITEM/RUN vocabulary), read the ground truth from RAM (dialog text names the species), let the decision distribution pick the action, translate choice tokens to tool calls, and log raw distribution + action per decision row. Tests must force a different choice and assert a DIFFERENT outcome \u2014 a test that passes on the hardcoded path is invalid. Keep recovery ladders as transport, not decision-makers.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "python-gameplay-de-hardcoding-jev-routing", "provider": "openrouter", "solved_at": "2026-09-24T21:58:50.836Z", "version": ""}