OpenRA.Bot is a Python-side RL/control package for OpenRA. It uses pythonnet to load the built game assemblies, calls the engine-side PythonAPI, and exposes a Gym-style environment for random agents, rule-based agents, and a baseline PPO training loop.
Status (2026-07-23): D0--D4 and the bounded post-D4 economy probes are complete, but no learned or spatial-economy candidate has passed promotion. The accepted build order remains the tick-3360 baseline. A completion-boundary harvester cap removed the extra-unit waste from local_density, but its fixed-window earned delta fell to 3625 versus first at 3650, proving the earlier 4300 result depended on a fourth harvester. The next iteration is engine-authoritative per-harvester mining-cycle telemetry and a fixed three-harvester diagnostic baseline, not a longer PPO run or combat work. Full architecture is documented in ARCHITECTURE.md.
- Python environment wrapper around the in-engine API
- Engine-side
PythonAPIbridge for local game start, stepping, state extraction, and action dispatch - Rule-based and random agents for smoke testing
- A custom
ActorCritic+PPOAgentbaseline with BC warm-start - Local game, local hosted lobby, and remote lobby connection helpers
- Build-order distance reward (DI-star / AlphaStar inspired)
- Asset-value reward with production-start / active-production credit
- Macro production action space for development-only training
- Decision-step policy gradient masking
- Action masking with
-infpenalty (prevents policy collapse) - Entity-based observations (Phase 1 —
observation_type="entity") - Headless local rollout and
SubprocVecEnvmulti-process training support
envs/openra_env.py: main Gym environment (MCV type detection fix, BO / asset reward, macro actions, decision mask support)envs/vector_env.py: spawnedSubprocVecEnvwrapper for multi-process headless rolloutenvs/wrappers.py:ShapedRewardWrapper(passthrough + diagnostics),AugmentedStateWrapper(frame stacking)utils/engine.py: loadsOpenRA.Game.dllandPythonAPIthroughpythonnetutils/obs.py: convertsPythonAPI.GetState()output into Python dictionariesutils/actions.py: encodes Python action dicts intoRLActionutils/net.py: local host / remote join / lobby helpersutils/PythonAPI.cs: engine bridge source (C#, ARM64 reflection bugs fixed —IsBuildingQueueOccupiedguard removed)utils/entity_obs.py: entity observation builder (14-dim per actor + 28-dim scalar, or 34 with goal conditioning)utils/goal_library.py: 4 build-order goals (economy/infantry/vehicle/balanced)agent/agent.py:RandomMoveAgent, state-awareRuleBasedAgent(cash/power/building checks),PPOAgentmodels/actor.py:VectorEncoder,SimpleEntityEncoder,MultiDiscretePolicy,ActorCriticwith GLU gating + multi-value-headsmodels/entity_encoder.py: MLP per-entity encoder + masked mean-pool (25K params)models/buffer.py: rollout buffer for PPO (GAE + dict obs support)scripts/train_rl.py: PPO training entry (entity/macro/headless/opponent/goal/teacher flags)scripts/warmstart.py: BC data collection + pre-training from RuleBasedAgentscripts/evaluate_policy.py: fixed independent evaluation for one checkpointscripts/select_checkpoint.py: independently ranks checkpoints; the development profile uses D0-bound completion, deadline, speed, and waste metricsscripts/search_build_order.py: bounded state-driven build-order search with successive halving and auditable D1 metricsscripts/create_selection_tie_fixture.py: creates an explicitly non-training current-environment tie fixture for selector integration checksreward_mode=development_efficiency: D0-bound development task withdevelopment_v1targets/deadline/economy observations and first-completion episodesscripts/remote_ppo.py: join a remote lobby and run trained model (with --stochastic --temperature)scripts/view_best.py: view best checkpoint in local game windowscripts/verify_asset_reward.py: reward headroom verificationARCHITECTURE.md: complete obs/action/reward/model specificationPLAN.md: long-term roadmap (AlphaStar-style architecture)REPORT.md: detailed experiment log and findings
The current execution path is:
envs/openra_env.pycallsutils/engine.pyto load the OpenRA assemblies.PythonAPI.StartLocalGame(...)or the lobby helpers initialize a match.PythonAPI.GetState()returns a simplifiedRLState.utils/obs.pyconverts that state into Python dicts.OpenRAEnvconverts the raw dict intofeature,vector, orimageobservations.- An agent chooses either a legacy dict action list or a
MultiDiscreteaction. utils/actions.pyandPythonAPI.SendActions(...)translate that into OpenRA orders.PythonAPI.Step()advances the simulation.
- A platform supported by your OpenRA build and
pythonnet - Python 3.8+
- A built OpenRA tree with
OpenRA.Game.dllandOpenRA.runtimeconfig.json - A mod and map that can be started from code, for example
ra
The Python package expects the compiled PythonAPI type to be available from OpenRA.Game.dll.
Recommended workflow:
- Keep the bridge source in
OpenRA.Bot/utils/PythonAPI.cs. - Run
scripts/sync_engine_bridge.ps1 -OpenRaRoot <OpenRA root> -Build. - Treat every file copied/patched in the outer OpenRA tree as an uncommitted
build artifact. Do not commit the bridge, compatibility patch, or an
OpenRA.Botgitlink to the outer repository.
sync_engine_bridge.ps1 is the tracked source of truth for integration. It
only copies utils/PythonAPI.cs; it never edits any other outer OpenRA source.
Engine upgrades and pythonnet compatibility must be solved inside this
repository rather than committed or patched into the outer repository.
The bridge bootstrap preloads the declared mod DLLs into the default .NET
context and seeds OpenRA's content-hash assembly cache. This preserves one
trait type identity under pythonnet without loading OpenRA.Mods.Common from
Python or editing outer engine sources. A bridge-only reset/build/smoke was
validated on 2026-07-24.
OpenRAApiBridge.cs is deprecated and should not be used by new code.
cd F:\Projects\OpenRA\OpenRA.Bot
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements.txtRule-based / manual smoke test:
python scripts/example_usage.pyBaseline PPO training:
python scripts/train_rl.pyCurrent recommended development-only PPO smoke launcher on Windows:
.\scripts\train_best.ps1 `
-Updates 10 -NumSteps 128 `
-WarmstartEpisodes 10 -WarmstartEpochs 20 `
-TeacherKlCoef 0.05 `
-RunName smoke_<name>Use a longer run only after the reachability/teacher-coverage gate passes. Useful overrides for an accepted experiment:
.\scripts\train_best.ps1 `
-Updates 150 `
-NumSteps 256 `
-MaxEpisodeTicks 1800 `
-WarmstartEpisodes 10 `
-WarmstartEpochs 15 `
-LearningRate 3e-5 `
-TargetKl 0.05 `
-EntCoef 0.005 `
-UpdateEpochs 4 `
-GoalAlignedWeight 0.6 `
-RunName obs_counts_goal_w06_stableThe launcher uses observation_type="entity", action_space_mode="macro", reward_mode="asset" through the default environment, headless mode, goal conditioning with -GoalAlignedWeight 0.6, and BC warm-start from RuleBasedAgent. With -TeacherKlCoef > 0, it also loads the BC action-type head so the teacher-KL penalty starts from the demonstrated distribution; use -NoLoadBcActionHead only for the explicit no-head ablation. Use -NoGoalConditioning for no-goal baselines.
After training, select a policy from numbered snapshots instead of trusting the shaped training-return checkpoint:
python -B scripts/select_checkpoint.py `
--run-dir checkpoints/<run-name> `
--seeds 101 --goal-indices 0,1,2,3Development promotion uses the accepted D0 baseline as an immutable deadline source. Its path and SHA are required and recorded in the selector output:
python -B scripts/select_checkpoint.py `
--run-dir checkpoints/<run-name> `
--selection-profile development_efficiency `
--development-baseline checkpoints/development_baseline_d0_20260717T102000Z_cost_corrected/development_baseline.json `
--development-baseline-sha256 1c9fcc514867a80ac49afb7fb709a83f3bdc409b276e401a576a2f004dc75830 `
--goal-indices 0,1,2,3The output is checkpoint_selection.json; it records all evaluated snapshots and the selected path. It intentionally does not overwrite model_best.pth, whose meaning remains “best rollout reward”.
To initialize a D2-compatible run, use the distinct environment contract (old checkpoints remain on the legacy scalar schema):
python -B scripts/train_rl.py `
--observation-type entity --action-space-mode macro --goal-conditioning `
--reward-mode development_efficiency --terminate-on-goal-completion `
--development-baseline checkpoints/development_baseline_d0_20260717T102000Z_cost_corrected/development_baseline.json `
--development-baseline-sha256 1c9fcc514867a80ac49afb7fb709a83f3bdc409b276e401a576a2f004dc75830This profile observes explicit current/target/remaining counts, deadline progress, and engine-authoritative cumulative earned/spent values. Its reward does not include gross asset, stored-resource growth, or active-production bonuses. Do not load a legacy checkpoint without an explicit architecture- drift experiment.
Run the D3 economy build-order search serially with a new output directory:
python -B scripts/search_build_order.py `
--goal-index 0 --candidate-budget 12 --finalists 4 --seed 101 `
--development-baseline checkpoints/development_baseline_d0_20260717T102000Z_cost_corrected/development_baseline.json `
--development-baseline-sha256 1c9fcc514867a80ac49afb7fb709a83f3bdc409b276e401a576a2f004dc75830 `
--output-dir checkpoints/<new-d3-search-directory>The script refuses to overwrite an existing directory, preserves exact target
and prerequisite quotas, and writes search_progress.json after every round.
The accepted run is
checkpoints/development_d3_economy_search_20260721T000000Z_retry2; its final
build_order_search.json SHA-256 is
b2b3836de160985b99deffb8bbc24cdb4b022868fe2f775f0664a544fb9198a1.
Create the goal-routed economy BC/DAgger initialization without PPO updates:
python -B scripts/train_rl.py `
--bin-dir F:/Projects/OpenRA/bin --headless `
--observation-type entity --action-space-mode macro `
--goal-conditioning --goal-indices 0 `
--reward-mode development_efficiency --terminate-on-goal-completion `
--development-baseline checkpoints/development_baseline_d0_20260717T102000Z_cost_corrected/development_baseline.json `
--development-baseline-sha256 1c9fcc514867a80ac49afb7fb709a83f3bdc409b276e401a576a2f004dc75830 `
--build-order-search checkpoints/development_d3_economy_search_20260721T000000Z_retry2/build_order_search.json `
--build-order-search-sha256 b2b3836de160985b99deffb8bbc24cdb4b022868fe2f775f0664a544fb9198a1 `
--warmstart-episodes 8 --warmstart-epochs 40 `
--dagger-episodes 4 --dagger-epochs 40 `
--max-episode-ticks 7921 --total-updates 0 --seed 101 `
--log-dir checkpoints/<new-bc-dagger-directory>The accepted handoff is
checkpoints/development_d3_economy_bc_dagger_20260721T000000Z.
Its model_0000.pth SHA-256 is
5f8b7352efe1b4138cb024f71f99e234314a8dc964fcc52cd4d32f514b54ab02;
an independent 8-launch greedy evaluation reproduced tick 3360 in every run.
D4 residual PPO is opt-in and requires an accepted external BC checkpoint:
python -B scripts/train_rl.py `
--observation-type entity --action-space-mode macro `
--goal-conditioning --goal-indices 0 `
--reward-mode development_efficiency `
--development-baseline <accepted-development-baseline.json> `
--development-baseline-sha256 <sha256> `
--warmstart-checkpoint <accepted-economy-bc.pth> `
--warmstart-checkpoint-sha256 <sha256> `
--residual-action-type-scale 0.5 `
--teacher-kl-coef 0.05 --bc-replay-path <accepted-bc-replay.pt>The original BC policy path is frozen; only the bounded residual action head
and critics train. model_0000 has an exactly zero residual. Use
select_checkpoint.py --selection-profile development_efficiency for every
run: its residual_promotion_gate requires at least 10 launches and a strict
3% median speed improvement without p90 or waste regression.
Post-D4 control-resolution screening is available through
scripts/evaluate_control_resolution.py. It binds both the D0 baseline and an
accepted D3 build-order artifact by SHA, journals each serial launch, and only
runs acceptance when a resolution is at least 3% faster without waste. The
10/5/2/1-tick economy matrix produced 3360/3320/3310/3295 completion ticks;
1-tick control was only 1.935% faster and roughly 7.3× slower in wall time, so
this direction was rejected. Fine-resolution testing also led to RuleBased
submitted-order and pending-placement commitments, preventing asynchronous
engine visibility gaps from creating duplicate production.
scripts/search_refinery_placement.py evaluates four deterministic refinery
cell rankings over bridge-authorized placement cells; it never exposes raw
(x,y) actions and leaves every other building on first-cell auto-placement.
All four economy strategies tied at completion tick 3360, so placement was not
promoted. However, authoritative cumulative earned at termination differed
(2000/2500/3000/2000), showing that the first-completion episode ends before
spatial-economy effects can be judged. The next evaluation must use a fixed
post-goal mining window before placement is considered for BC/PPO.
scripts/evaluate_post_goal_mining.py implements that separate evaluation
contract. After first completion it freezes all teacher actions for a fixed
engine-tick window and measures authoritative cumulative-earned delta; this
does not change D2 reward or training termination. In the 1800-tick screening,
local_density earned 4300 versus first at 3650, but ended with 1100 surplus
from production already in flight at completion. The no-waste paired gate
therefore rejected it. The next task is completion-boundary queue scheduling,
not adding income reward or placement actions to PPO.
scripts/search_completion_queue_schedule.py implements that scheduling gate.
RuleBased can apply explicit per-item production caps and the evaluator records
completion/final composition plus post-goal deltas. With harv=1, all four
placement strategies retained tick 3360 and exactly three final harvesters,
but fixed-window earned was 3650/2500/3625/3500 for first/nearest/local-density/
distance-weighted. No candidate met the +10% threshold, so paired acceptance
was not launched. Placement is therefore still excluded from BC/PPO. The next
gate is per-harvester pickup/delivery/load/travel/idle telemetry, followed by a
fixed-composition diagnostic before any harvest-control action is added.
The first fixed-composition diagnostic is complete in
checkpoints/development_harvester_telemetry_20260723T020000Z: 8/8 launches
completed at tick 3360, ran through tick 5160, retained three harvesters, and
earned 3650. The evaluator records movement/idle observations per harvester,
but the current bridge has no cargo, pickup, delivery, or assignment events;
these fields are intentionally not inferred from position. Do not promote a
harvest action or change reward until the bridge exposes those authoritative
cycle signals.
The bridge now exposes those signals. It reports harvester cargo/capacity,
activity phase (harvesting, returning, unloading), resource target cell,
and reserved refinery. Cargo edges provide authoritative pickup/delivery and
complete-cycle counts. In the formal 8-launch baseline, all episodes preserved
3360→5160 and earned 3650; the three harvesters completed 6/3/2 cycles per
episode with mean durations 386.7/436.7/400 ticks. The next experiment is a
bounded RuleBased resource-assignment screen, not PPO or a reward change.
Parallel rollout can be enabled from train_rl.py:
python scripts/train_rl.py --observation-type entity --action-space-mode macro --headless --num-envs 4Remote rule-based control:
python scripts/remote_rule_based.py --host 127.0.0.1 --port 1234 --slot Multi0Remote PPO action debugging:
python scripts/remote_ppo.py --host 127.0.0.1 --port 1234 --slot Multi0Remote PPO training:
python scripts/train_rl.py --remote-host 127.0.0.1 --remote-port 1234 --remote-slot Multi0from envs.openra_env import make_env
env = make_env(
bin_dir="F:/Projects/OpenRA/bin",
mod_id="ra",
map_uid="b53e25e007666442dbf62b87eec7bfbe8160ef3f",
ticks_per_step=10,
observation_type="vector",
enable_actions=["noop", "move", "attack", "produce", "build", "deploy"],
)
obs, info = env.reset()
for _ in range(1000):
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
break
env.close()OpenRAEnv currently supports four observation modes:
feature: returns the raw Python dict built fromPythonAPI.GetState()vector: flattened numeric observation for MLP policiesimage:128 x 128 x 10semantic map for CNN-style policiesentity: fixed-cap entity tensor plus scalar features for the current lightweight entity encoder
This is the most complete mode and is the best option for debugging. It includes:
actorsresourcesproductionproducible_catalogplaceable_areascashresources_totalresource_capacitypowermy_owner
Each actor now exposes two order-related fields:
available_orders: a filtered list intended for bot/RL logicavailable_order_ids: the raw order ids exposed by engine traits
Important nuance: available_order_ids is closer to "what traits are present on this actor", while available_orders is the safer field to use for decision logic. For example, transformable buildings may expose a raw Move order id even when they should not be treated as currently mobile.
Current vector layout:
- Up to 100 friendly units, 6 features each
- Up to 100 enemy units, 5 features each
- 7 resource/power slots
- 2 map-size slots
Important caveat: the resource/power section is currently placeholder-filled in envs/openra_env.py, so the vector observation does not yet fully use the economy state already available from PythonAPI.GetState().
Current image layout:
- Shape:
(128, 128, 10) - Channels currently used reliably:
- friendly infantry / non-infantry
- enemy infantry / non-infantry
- Channels for resources, cash, and power are not fully populated yet
This is the current default for new PPO experiments. It uses utils/entity_obs.py to build:
entities: up toMAX_ENTITIESactor rows with per-actor featuresentity_mask: valid-row maskscalar: compact economy / power / queue / count / game-state features
The lightweight entity encoder is enough to clone the current RuleBasedAgent. Goal conditioning is the strongest current training signal; after adding count-aware scalar features, the remaining issue is PPO stability rather than pure observability.
The default multidiscrete RL action space is:
[action_type, unit_idx, target_x, target_y, target_idx, unit_type_idx]
Supported action types depend on enable_actions, but the usual set is:
noopmoveattackproducebuilddeploy
Semantics:
move: usesunit_idx,target_x,target_yattack: usesunit_idx,target_idxproduce: uses queue actor index +unit_type_idxbuild: uses queue actor index +unit_type_idx+ target celldeploy: usesunit_idx
The environment also accepts legacy Python dict actions, which is what the rule-based agent uses.
For development-only training, action_space_mode="macro" collapses the meaningful decision to the action_type head:
noop, produce:powr, produce:proc, produce:barr, produce:weap, produce:e1, ...
The remaining argument heads are masked to index 0. MCV deployment and finished-building placement are handled by environment automation. This was added because the original 6-head action space made successful production a low-probability joint event (produce + correct queue + correct unit_type), which was too hard to reinforce reliably in single-env PPO.
info["action_mask"] currently includes some of the following fields:
action_typemove_maskattack_maskdeploy_maskproduce_queue_maskproduce_unit_type_maskbuild_maskbuild_unit_type_maskunit_idxtarget_idxtarget_xtarget_yunit_type
These masks are no longer only heuristic action-type hints. The current implementation mixes:
- engine-side feasibility checks for
move,attack, anddeploy - queue-state- and placement-driven checks for
produceandbuild - per-head masks consumed by
PPOAgentduring both sampling and training
Current behavior:
move_mask: only set when the actor has a feasible move in a nearby neighborhoodattack_mask: per-attacker / per-target feasibility matrixdeploy_mask: checked through engine feasibilityproduce_queue_mask: only queues that are enabled, empty, and can actually produce something in the current catalogbuild_mask: only queues with a completed item and a currently available placement areatarget_x/target_y: conditioned on the selected actor or queue- move targets are restricted to a local neighborhood around the selected actor
- build targets are restricted to coordinates present in
placeable_areas
Remaining limitation: target_x and target_y are still masked independently rather than as a joint (x, y) cell distribution, so some invalid coordinate pairs can still be sampled.
The default reward in envs/openra_env.py is development-oriented, not combat-oriented. The current default reward_mode="asset" rewards:
- cost-weighted growth in owned actors (
AssetValueTracker) - starting new production items
- keeping production active
- canceling in-progress production, as a penalty
- optional idle-cash and power-deficit penalties, currently disabled by default because per-step penalties can swamp one-time asset gains
reward_mode="legacy" keeps the older build-order / unit-count shaping path. The asset reward fixes the capped [powr, proc, barr] build-order reward and makes army production visible, but it is still not enough on its own for strong tactical play or decisive improvement over a scripted macro teacher.
OpenRAEnv.reset() supports three startup modes:
- local single-player start through
PythonAPI.StartLocalGame(...) - host-local lobby flow through
env.configure_host(...) - remote server join through
env.configure_remote(...)
See utils/net.py for the exact lobby helper flow.
For remote control, the current flow is:
env.configure_remote(...)reset()joins the server- the client claims a slot, acknowledges the selected map, and marks itself ready
- lobby/network state is pumped until the host starts the game
- once the world exists, normal observation / action stepping begins
Recent bridge changes were specifically made to keep network traffic progressing while still in the lobby, so remote clients can stay synchronized through the lobby-to-game transition.
Policy collapse (KL explosion to 15+)Fixed with-infmask penaltyMCV deploy reward=0Fixed with type-based_is_buildingdetection (ARM64 reflection bug workaround)Reward signal too sparse (0.7% non-zero)Fixed with BO distance reward + producing_per_step (now 98%+ non-zero)Action mask penalty insufficientFixed:log(clamp(1e-6))→-infShort rollout withFixed innum_steps < seq_lenproduced zero PPO batchesmodels/buffer.pyMacro action auto-placement could fail silentlyFixed: done buildings are placed by queue actor idSingle-process-only rolloutImproved with headlessSubprocVecEnvsupportTeacher-KL could produce NaN underFixed: reduce only over legal decision/action pairs-infmasks or rollout paddingPPO long-sequence new log-probs were aligned to the wrong timestepFixed: flatten(batch, env, time, action)without moving the action axisMacro masks could permit visible but locked technologyFixed: bridge exposesBuildable; Python masks, teacher, and macro resolver require itDevelopment selection rewarded gross production and copied deadlines into codeFixed: the D0-bound profile validates baseline SHA/DLL provenance, ranks completion/deadline/speed/waste, and preserves earlier snapshots on exact ties
- PPO can currently recover, but not exceed, the BC/RuleBased macro baseline under independent evaluation
- The vehicle-only BC-policy-feature-freeze ablation preserves BC completion through PPO update 6 (versus roughly update 2 for replay CE), but has no faster completion or other strict advantage; do not promote tied snapshots or tune more regularizer coefficients on this solved development-only task.
model_best.pthselects shaped rollout return; usescripts/select_checkpoint.pyto choose the policy checkpoint- Teacher-KL starts from the BC action head by default, but its current teacher does not cover the full vehicle-tech path within the 1800-tick demo horizon
- The current expert prior is a hand-written macro script without scouting, combat, tech transitions, or opponent adaptation
- All four current goals have 0 completion in the 1800-tick evaluation; validate reachability and horizon before interpreting PPO comparisons
- Goal conditioning is implemented and is the current strongest improvement; staged / moving-goal variants underperformed fixed goal weights
- Rewards are still development-heavy; combat and terminal win/loss are still too thin in the main training loop
- The bridge reports authoritative game-over and local win-state, so
OpenRAEnvcan distinguish terminal win/loss from a time-limit truncation. A built-innormalbot smoke confirms opponent startup, but combat PPO is still gated on a non-catastrophic scripted fixed-opponent baseline. macro_combatis a separate checkpoint schema that retains production categories and adds feasibility-gatedattack:visible,retreat:base, andrally:base. The bridge reports loss values fromPlayerStatistics.DeathsCost; opponent reward uses incrementalenemy_loss − own_loss, not fog disappearance. The most reliable post-goal/minimum-five vehicle baseline (14400 ticks, seed 101) completed at tick 7240 but finished at potential0.782with own loss11250, enemy loss8650, and trade-2600. A two-seed base-rally/mixed-force experiment was worse: both games ended inLost(mean trade-22750) andrally:basewas never selected because new units already spawn near the Construction Yard. Combat BC/PPO remains blocked; the next viable control hypothesis needs a meaningful forward/defensive spatial objective, not another base-relative macro.PythonAPI.GetState()is expensive (scans all world state every call, heavy reflection usage)target_x/target_ymasking is still factorized rather than fully cell-jointPythonAPI.csuses reflection to accessOpenRA.Mods.Commontypes — fragile across OpenRA versions and platforms- Multi-environment rollout works in headless mode, but Windows multiprocessing / CLR process management remains operationally fragile and needs more soak testing
- Deploy mask blocks building undeploy (
openra_env.py): The deploy action mask excludes known building types (fact,afld,weap, etc.) from undeploying via DeployTransform. This prevents the agent from constantly undeploying the Construction Yard and canceling in-progress production. In a real game, undeploying to relocate the base is a valid strategy. TODO: Remove the building-type blocklist and let the agent learn the cost of interrupting production via reward penalties (e.g. production-cancel penalty, time-waste penalty). - Production queue safety (
openra_env.py,PythonAPI.cs): a queue with any active orDonebuilding item stays occupied until placement is accepted. Python masks/teacher/macro resolution enforce this policy, and the synchronized C# bridge guard rejects directStartProductionbypasses as defense in depth. - Per-category produce mask (
openra_env.py): Theproduce_unit_type_maskblocks unit types whose production queue category (e.g. "building") is already occupied. This is the soft counterpart of the C#-side guard above. TODO: Re-evaluate once the agent can reliably complete the produce→build cycle.
- Use
featureobservations first when debugging action execution. - For remote debugging, start with
scripts/remote_rule_based.pyorscripts/remote_ppo_debug.pybefore running long PPO training jobs. - Treat the current PPO stack as a baseline to iterate on, not a final trainer.
- For current PPO experiments, prefer
scripts/train_best.ps1defaults: count-aware entity observations, macro actions, asset reward, goal conditioning w=0.6, and headless. Supplying-TeacherKlCoefautomatically keeps the BC action head unless-NoLoadBcActionHeadexplicitly requests the ablation. - Track
mean_reward,last20, KL, entropy, batches, decision steps, action distribution, and reward components; use independent goal metrics—not reward alone—for checkpoint selection. - Evaluate numbered snapshots before promotion. Development runs must use
--selection-profile development_efficiencywith the accepted D0 baseline path and SHA. It ranks completion reliability, deadline success, speed, potential, and lower waste; gross production is telemetry only andmodel_best.pthremains a training-return diagnostic. - If policy quality is the goal, first run the goal-reachability/teacher-coverage matrix, then build goal-specific demonstrations; do not scale model size or PPO updates before that gate passes.
- If sample efficiency is the goal, optimize state extraction and keep hardening multi-env headless rollout.
- Local start fails: check
bin_dir,mod_id, andmap_uid, and confirm the OpenRA build artifacts exist. - Python cannot load the engine: verify
OpenRA.runtimeconfig.jsonand the required assemblies are present inbin_dir. - Remote join enters the lobby but does not stay synchronized: rebuild OpenRA after syncing
utils/PythonAPI.cs, because remote-lobby behavior depends on the latest bridge code. - Production/build actions appear invalid: inspect
production,placeable_areas,available_orders, andavailable_order_idsfromfeatureobservations first. - Actor indices behave strangely: remember that
unit_idxandtarget_idxare mapped through the latest cached unit-id lists, not raw actor IDs. - A transformable building appears to have
Move: checkavailable_order_idsvsavailable_orders. The raw field may still include transform-related move orders, while the filtered field is the one intended for control logic.
This section documents a deep analysis of the current architecture, focusing on environment interface issues and the gaps between the current baseline and an AlphaStar-style RTS RL agent.
- Complete closed loop: pythonnet bridge → state extraction → Gym env → PPO training all wired up end-to-end
- Rich action masking: engine-side feasibility checks for move/attack/deploy, queue-state-aware produce/build masks
- Multiple connection modes: local game, host-local lobby, remote server join
- Auto-placement: MCV auto-deploy and Done-item auto-place reduce action complexity
The current observation_type="vector" uses a fixed 100-slot bin for both friendly and enemy units (openra_env.py:1276-1304). This has several critical flaws:
- Fixed capacity truncation: RTS unit counts vary from 1 (start) to 50+ (late game). Fixed 100 slots waste space early and may overflow late.
- No relational information: Units are encoded independently — the network cannot learn which units are fighting which.
- No terrain awareness: The vector obs lacks terrain type, passability, and fog-of-war boundary information.
- Order-dependent encoding: Unit order in the fixed slots changes across ticks, making it hard for the network to track identities.
AlphaStar's approach: Entity-based attention (Transformer over variable-length entity list) + spatial grid encoder (ResNet over minimap), which naturally handles variable entity counts and preserves spatial/relational structure.
The current action space is MultiDiscrete([action_type, unit_idx, target_x, target_y, target_idx, unit_type]) — 6 independently-sampled categorical heads.
target_x/target_yare independently factorized (openra_env.py:1167-1173): Valid x-coordinates can be paired with invalid y-coordinates, producing impossible targets. This should be a joint 2D spatial distribution.- No hierarchical conditioning: All 6 heads are predicted from the same shared feature vector without autoregressive conditioning. The network cannot learn that
targetdepends onunit_selectionandaction_type. unit_idxis a fixed categorical: 100-way classification over slots. Unit ordering changes across ticks, so index 5 means different units at different times.- Missing action types: No
harvest,repair,sell,guard,stop, orcancel_productionactions.
AlphaStar's approach: Hierarchical autoregressive action space (action_type → unit_selection → target) with pointer networks for entity selection and 2D spatial logits maps for spatial targets.
PythonAPI.GetState() (PythonAPI.cs:459-651) rebuilds the entire RLState every call:
- Heavy reflection usage: Economy traits (
PlayerResources,PowerManager), production queues, and building placement are all accessed via reflection. EachGetState()call repeatsAssembly.GetType()scans andMethodInfolookups. - Full map scan for placeable areas:
CollectPlaceableAreas()iterates every cell on the map (map.AllCells) and checksCanPlaceBuilding+IsCloseEnoughToBasefor each Done item type. This is O(cells × building_types). - No incremental updates: Even if only one tick passed and nothing changed, the entire state is rebuild from scratch.
Target: Cache reflection calls in static fields, use incremental state updates, and optimize spatial queries.
SubprocVecEnv now supports spawned headless workers, so the old single-environment hard limit is gone. This improves wall-clock sample throughput, but recent 4-env and 8-env experiments still did not produce a decisive improvement over the scripted teacher. The remaining issue is algorithmic and task-level: on-policy PPO still has high variance on long-horizon RTS development, especially when the reward is development-only and the teacher already covers the easy macro path.
Target: Keep hardening 8-64 parallel environments, but pair scale with stronger priors, teacher-KL, goal conditioning, and combat / win-loss rewards.
The current reward (openra_env.py:407-467) only incentivizes:
- Unit/building count increase
- Starting new production
- MCV deployment
Critically missing:
- Combat rewards: Damage dealt, enemy units killed
- Win/loss signal: The most important reward in any competitive game
- Resource efficiency: Ore collection rate, not just total
- Map control / exploration: Reconnaissance value
| Issue | Location | Impact |
|---|---|---|
| Reflection not cached | CollectProductionInfo(), CollectPlaceableAreas() |
~50-200ms per GetState() call |
| Full-cell-scan for build areas | CollectPlaceableAreas() map.AllCells loop |
O(cells × types) per call |
| Building queue single-item guard | SendActions() IsBuildingQueueOccupied |
Prevents advanced queue management |
| Deploy mask blocklist | openra_env.py _building_types hardcoded set |
Prevents agent from learning base relocation |
| Dimension | Current OpenRA-Bot | AlphaStar |
|---|---|---|
| Entity encoding | Fixed 100-slot bin, order-dependent | Transformer over variable-length entity list |
| Spatial encoding | 128×128×10 semantic map (channels mostly empty) | ResNet over rich minimap + camera view |
| Action structure | Independent 6-way MultiDiscrete categoricals | Hierarchical autoregressive with pointer networks + 2D spatial heads |
| Unit selection | Fixed categorical over 100 slots | Attention-based pointer network over entity embeddings |
| Spatial target | Independent x, y categoricals (validity broken) | 2D logits map with spatial masking |
| Network core | MLP-based encoder + optional LSTM | Deep LSTM + Transformer + ResNet |
| Training parallelism | Headless SubprocVecEnv works at small scale; needs hardening and better sample efficiency |
Thousands of parallel environments (TPU pods) |
| Opponents | Single built-in bot | Self-play league with PFSP, historical agents, exploiter agents |
| Reward | Development shaping only | Win/loss + game statistics |
| Pre-training | None (random init) | Supervised pre-training on human/expert replays |
| Curriculum | Fixed map | Progressive difficulty + map diversity |
See PLAN.md for the full implementation plan. Below is a high-level summary:
- Fix
target_x/target_yindependent factorization → joint 2D mask - Complete
unit_typesmapping from CSV or game data - Add combat + win/loss rewards
- Establish performance baselines (random, rule-based, PPO, built-in bot)
- Entity-based observation builder (variable-length, feature-rich)
- Extended spatial observation with terrain, fog, threat channels
- Scalar observation (economy, power, game time)
- Dict-based
gym.spaces.Dictobservation space - Backward compatible with
observation_type="vector"
- Hierarchical autoregressive action structure
- Spatial action head (2D logits map, not factorized x/y)
- Pointer network for unit selection (attention-based)
- Joint 2D spatial action masks
- Expanded action types (harvest, repair, sell, guard, stop)
- Entity Transformer encoder (self-attention over variable-length entities)
- Spatial ResNet encoder (deeper, richer than current CNN)
- AlphaStarActorCritic: integrated encoder + LSTM + hierarchical head
- Maintain backward compatibility with old
ActorCritic
SubprocVecEnvfor parallel environments (target: 8-64 envs)- Self-play environment (two Python-controlled players)
- Elo rating evaluation framework
- 5x+ throughput improvement over single-env
- League training with PFSP (Prioritized Fictitious Self-Play)
- Modular reward system (win/loss, combat, economy, exploration)
- Curriculum learning (economy-only → static opponent → full combat → map generalization)
- Supervised pre-training via behavior cloning from built-in bots
- Cache reflection calls in static fields → ~10x speedup for GetState()
- Incremental state updates (
GetStateDelta()) - Batch order feasibility checks
- Target: GetState() < 10ms (currently 50-200ms)
-
Parallelism strategy: Python multiprocessing with
spawnstart method (each subprocess loads its own CLR) vs. single-process multi-instance (requires engine-side changes) -
Transformer scale: Start with 3 layers / 4 heads / 256-dim model. Scale up only after training stability is proven.
-
Action space granularity: Start with 8-10 action types. Add more only when the agent masters the basics.
-
Curriculum vs. end-to-end: Curriculum (Phase 5) is recommended for training stability, but the architecture should support end-to-end from day one.
-
Self-play vs. fixed opponents: Start with built-in bots for baseline, then add self-play gradually. Full league training is the last piece.
MIT