Files
OpenMaidEngine/docs/engine-re.md
2026-07-10 21:37:42 -04:00

975 lines
78 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Native-engine reverse engineering (Ghidra + MCP)
Static RE of the **unpacked** `AGE.EXE` engine image, driving Ghidra 12.1.2 via the
bethington/ghidra-mcp bridge. This is the home for decompiled native-op findings — the class of logic
the scripts call but that lives compiled in the engine (decision→scene, call-script dispatch, op 0x60,
the gfx command-buffer). Opcode semantics recovered here also flow into `vm-map/opcodes.toml`.
Related: `docs/scjump-progression.md` (the SCJUMP decoder that hit this wall), `name-resolution.md §1`
(call-script), `vm-mapping-plan.md` appendix (why the exe is packed + the runtime-dump route).
---
## Runbook — the Ghidra + MCP loop
**One-time setup (done 2026-07-07):**
- **MCP server:** bethington/ghidra-mcp, cloned to `S:\Game Hacking\ghidra-mcp`. We used the **prebuilt
extension** `GhidraMCP-5.14.2.zip` (installed in Ghidra via File > Install Extensions) — this skips
the Maven/Java-21 build. The Python **bridge** runs from a venv (`.venv`, Python 3.11, `pip install .`);
no `uv` needed. Registered in Claude Code via `.mcp.json` at the workspace root:
`{"mcpServers":{"ghidra":{"command":"S:\\Game Hacking\\ghidra-mcp\\.venv\\Scripts\\bridge-mcp-ghidra.exe","args":["--transport","stdio"]}}}`.
- **In Ghidra:** enable the GhidraMCP plugin (File > Configure) and **Tools > GhidraMCP > Start MCP
Server** (serves `http://127.0.0.1:8089/`). The bridge talks to that; Claude reaches the bridge over stdio.
**Loading the engine image (IMPORTANT — the language gotcha):**
- Import `age-reimpl/build/engine-dump/range_00400000.bin` (the module dump: 2,490,368 bytes, the full
0x400000 module image; VA→file offset = `VA 0x400000`).
- **Format = Raw Binary, Language = `x86:LE:32:default`, Image Base = `0x400000`.** Ghidra's language
picker offers `x86:LE:32:System Management Mode` as the "closest" match — **do NOT use it.** SMM is a
16-bit *segmented* (segment:offset) variant for BIOS/SMRAM; it mis-decodes flat 32-bit code (it loaded
with addresses like `0000:0000`/`0025:ffff` and produced **0 functions**). The plain `default` variant
is correct and yielded **2,721 functions**.
- We drove the (re)import over MCP: `import_file(language="x86:LE:32:default", compiler_spec="windows",
auto_analyze=false)` → `set_image_base(0x400000)` **before** analysis (so absolute-address refs resolve)
→ `run_analysis`.
- **Load sanity check (AGF-decoder landmark):** at VA `0x474f23`, `CMP word ptr [ESI + 0x4], 0x4d42`
(the `BM`/BMP-magic check) confirms the image is correctly based + decoded.
- IAT reconstruction — **tried, DOESN'T WORK on this binary (2026-07-09):** `bin/pe-sieve32.exe /pid
<PID> /imp 3 /dmode 3 /dir build/pe-sieve` (run from **PowerShell**, not Git Bash — it mangles
`/flags`) ran fine but the game is packed with a **zeroed IAT** resolved via `GetProcAddress` at load,
so there is no conventional import table to rebuild. Of 363 "imports" it emitted, only ~17 are genuine
(`in_main:1`): the packer bootstrap (`LoadLibraryA`/`GetProcAddress`/`GetModuleHandleA`/`VirtualAlloc`/
`VirtualFree`) + a one-per-DLL seed block at RVA `0x202bfc` (`d3d9.Direct3DCreate9`,
`user32.RegisterClassExA`, `gdi32.GetStockObject`, `winmm.timeSetEvent`, `advapi32.RegOpenKeyA`,
`shell32.SHGetSpecialFolderPathA`, `oleaut32.Variant*`, `kernel32.RaiseException`, …). The other 300+
are stray pointer-shaped DWORDs mis-resolved to "first export at module base" (e.g. `msvcrt._wstrtime_s`
30×, `in_main:0`, non-terminated). ⇒ **do not graft pe-sieve output** — grafting the noise would inject
wrong import names. The game's hot APIs (`ReadFile`/`CreateFileA`/`timeGetTime`/d3d9 device methods) are
`GetProcAddress`-resolved into private pointer tables, invisible to a static IAT scan. Report/dump left
at `build/pe-sieve/process_<pid>/` (disposable).
**→ The Frida import-map approach — ✅ DONE 2026-07-09 (replaced pe-sieve).** Named the
dynamically-resolved APIs at their call sites via the LIVE process. `tools/frida/map_imports.py`
(read-only, plain-JS): (1) Frida-reads all loaded modules' export tables → `{runtime_addr → dll!Func}`
(23,342 exports); (2) scans the `0x400000` module for aligned DWORDs holding those addresses → `RVA →
name` (ASLR-stable: RVAs into the fixed main module transfer to the dump even though the DLL targets
relocate); (3) a `run_script_inline` pass labels the `/v2` image `imp_<dll>_<func>`. **Result: the packer's
rebuilt core IAT lives at RVA `0x16f000` (VA `0x56f000`) — 248 imports labeled** (kernel32 129, user32 57,
winmm 20, gdi32 17, advapi32/ole/oleaut/version/ntdll), 0 clobbers. **Validated:** `FUN_0044f390` now reads
`(*imp_kernel32_CreateFileA)` / `(*imp_kernel32_SetFilePointer)` at its resolver I/O; `sleep_timer_arm` reads
`(*imp_winmm_timeGetTime)()` — pinning the long-standing `DAT_0056f3d4` = **timeGetTime**. Tool:
`tools/frida/map_imports.py [--recon]` → `build/import-map.json`; plan
`docs/superpowers/plans/2026-07-09-frida-import-map.md`.
**Known limit (by design):** only the module-resident IAT is labelable. `d3d9`/`shell32`/`dsound`/CRT are
`GetProcAddress`-resolved into HEAP (not in the `0x400000` dump), so they aren't labeled — and D3D9 is used
via COM vtables (`Present` = device vtable slot 17, see `probe_present.py`), not an import thunk, so this
costs us nothing on the render path. 29 isolated singleton matches were set aside (`build/import-map-singletons.json`),
not auto-applied.
---
## Master key — the opcode→handler dispatch table (2026-07-07, anchored)
The interpreter dispatches each op via a per-context handler table, **fully anchored**:
> **`handler(op) = ctx[0x26c93 + op]`** (word index) **= `*(ctx + 0x9b24c + op*4)`**
> — `ctx` = the engine context (`esi` in handlers, thiscall; `param_1` in the decompile of the
> registration routine).
The registration routine **`FUN_00413860`** first fills `0x400` (1024) slots starting at
`ctx[0x26c93]` with a **default handler `FUN_004162b0`** (op 0's slot), then overrides specific
opcodes: `ctx[0x26c93 + op] = <handler_va>`. So **opcode = (word_index 0x26c93)**. Cross-check:
`ctx[0x26e3f] = 0x427fb0` (byte offset `0x9b8fc`) → op `0x26e3f 0x26c93 = 0x1ac`.
**Why this matters:** the Kelebek `u00XXXXXX` opcode names encode handler VAs from *Kelebek's* build,
which **drift** in ours. This table resolves the *real* handler for any opcode in our image — the
general fix for VA drift project-wide. To find op `N`'s handler: read `ctx[0x26c93 + N]` from the
`FUN_00413860` decompile (or `*(ctx + 0x9b24c + N*4)` at runtime).
### Materialized + applied image-wide (2026-07-09)
The table is no longer resolved op-by-op by hand — it is **extracted once and applied to the whole
image**. `tools/ghidra_handler_map.py` parses the override stores in `FUN_00413860` (dump at
`build/engine-dump/FUN_00413860.disasm.txt`) → **`build/op-handler-map.json`** (`{op → handler VA}`,
**420 overrides**). Regenerate: `py -3.11 -X utf8 tools/ghidra_handler_map.py
build/engine-dump/FUN_00413860.disasm.txt --check`. The `--check` diffs the derived handlers against the
handler VAs mentioned in `vm-map/opcodes.toml` prose and found **0 real drift** — the only 7 flags are
ops whose toml text records the *worker* VA, not the handler (`0x20c→0x4174a0`, and the `0x21c0x243`
cluster entries), each already matching the recon tables below.
A one-shot Ghidra script (via `run_script_inline`; needs `GHIDRA_MCP_ALLOW_SCRIPTS=1`) then labeled the
image from that map: **281 raw `FUN_`/`LAB_` handlers renamed `op_0xNN_handler`, 107 bare handler VAs
turned into functions, 31 hand-named handlers preserved** (source `USER_DEFINED` is never renamed), and
a plate comment `opcode 0xNN dispatch handler; ctx[0x26c93+op] in FUN_00413860` set on every one
(appended to existing decode comments, never clobbering). The one shared handler `0x416650` (ops
`0xaf`/`0x1a8`) is `op_0xaf_0x1a8_handler`. ⇒ every dispatch handler in the image now self-identifies its
opcode; a bare `op_0xNN_handler` is a handler not yet role-RE'd. Enrich with a descriptive name +
decode when you reverse one (the generic name is a floor, not a final).
> **⚠ Two-program gotcha (cost time 2026-07-09).** The Ghidra project holds **two** imports named
> `range_00400000.bin`: the GOOD one at project path **`/v2/range_00400000.bin`** (`x86:LE:32:default`,
> image base `0x400000`, 4308 functions — all our annotations live here) and a BROKEN early import at
> **`/range_00400000.bin`** (the `x86:LE:32:System Management Mode` mis-import: base `0000:0000`, **0
> functions**; see the language gotcha in the runbook). After a Ghidra restart the broken one can become
> active. **Always confirm `get_current_program_info` shows base `0x400000` / 4308 functions (or
> `switch_program /v2/range_00400000.bin`) before doing anything** — `run_script_inline` runs against the
> GUI's active program, so a wrong-program script would mutate/measure garbage.
**Other confirmed engine-context offsets** (`ctx`/`esi`): `+0x53d14` = current gfx-object index;
`+0x53d88` = per-object cmd-type table (stride `0x78` = 120 bytes); operand-fetch helper =
**`vm_operand_fetch`@`0x41b940`** (thiscall, `ecx=ctx`, arg = operand index → returns the operand value);
**`vm_operand_write`@`0x425fb0`** = the counterpart store; **`vm_operand_lvalue`@`0x415f30`** = the
companion index/pointer accessor.
**Hot-helper naming pass (2026-07-09, lever #2).** Ghidra's Function ID analyzer names **0** functions on
this image (the bundled FidDbs don't cover the VC9/VS2008 static runtime; ~3,660 of 4,428 funcs stay
`FUN_`), and the library workers we actually touch were already hand-named (`gfx_object_query_source_slot`,
etc.). So "STL/CRT auto-naming" had little to add — but the recon (rank unnamed funcs by call-count)
surfaced the real win: ubiquitous **documented-but-unnamed helpers**. Named the top 5 (~2,400 call
sites): `vm_operand_fetch`@`0x41b940` (1021 refs), `vm_operand_write`@`0x425fb0` (188),
`vm_operand_lvalue`@`0x415f30`, plus two CRT primitives identified by behavior —
`__security_check_cookie`@`0x54f981` (692; compares `__security_cookie`=`DAT_005c28c0`) and
`operator_new`@`0x5502be` (533; `_malloc`+`__callnewh`+throw `bad_alloc`). ⇒ every handler now reads
e.g. `vm_operand_fetch(2)` not `FUN_0041b940(2)`. No FidDb generation (out of scope, low ROI). Rename
hot unnamed funcs by call-count when de-noising further; there is no registry file for these — the
Ghidra name is the record.
**These `ctx` offsets are now a typed struct (2026-07-09).** The canonical field map is
`vm-map/engine-ctx.toml` → generated `docs/engine-ctx-reference.md`; a `run_script_inline` pass created
an `EngineCtx` Ghidra struct and retyped **all 419 dispatch handlers' `this` to `EngineCtx *`**, so they
decompile `ctx->cur_ctx_index` / `ctx->cmd_type_table` / `ctx->run_state_flags` instead of `param_1 + 0x…`
(verified: `sleep_op_0xc8`, `gfx_op_0x215_query_source_slot`). Add a field: edit `engine-ctx.toml`, run
`engine_ctx_build.py --build`, re-apply the struct. (The VM global bank `G[…]` is separate — `globals.toml`.)
---
## Findings
### op `0x1a2` (`u00428010`) is a GRAPHICS command-buffer op — NOT save, NOT decision→scene (2026-07-07)
The SCJUMP slice assumed `u00428010` resolved a decision value to a scene. **That premise is wrong**,
and pinning the *real* handler via the dispatch table above corrects two layers of confusion:
- **VA-drift trap:** Kelebek's `u00428010` = op `0x1a2`. But Kelebek's raw VA `0x428010`, in *our*
build, sits inside a *different* handler `0x427fb0`, which is **op `0x1ac`** (per the table:
`ctx[0x26e3f]=0x427fb0`). Op `0x1ac` is a **save-path** op — its handler formats
`%s\SAVE%2.2d.DAT` (format string `0x571e70`) and is multi-operand. Reading the raw VA gave the
wrong opcode.
- **Op `0x1a2`'s real handler = `FUN_0042d360`** (`= ctx[0x26c93+0x1a2] = ctx[0x26e35]`), argc 1. It:
sets the **current gfx-object cmd-type to 3** (`*(ctx+0x53d88 + ctx[0x53d14]*0x78) = 3`), fetches
operand 1, formats a key with `"%c%8.8x"` (format string `0x5714e0`) of `(3, operand)`, and calls
`FUN_0042cf70(key, &operand)`. This is a **graphics command-buffer registration op**, not save and
not scene-load.
- **Consequence — the decision→scene premise is discredited.** The FIELD snippet
`lookup(0x5f0ed, 0x62ccf); mov(ptr,1); lookup(0x5f0ed, 0x62ccf); u00428010(ptr)` (next op `0x21b`,
also gfx-family) is a **graphics/UI operation**, not scene sequencing. So `u00428010` does **not**
resolve decision→scene. **The real decision→scene mechanism is unidentified** — it belongs with the
call-script / script-load dispatch (`name-resolution.md §1`), the next target for this loop (now
armed with the dispatch table to resolve the call-script handler directly).
**Lesson:** never analyze a native op by its Kelebek `u00XXXXXX` VA directly — always resolve the real
handler through the dispatch table (`ctx[0x26c93 + op]`). The raw VA is off by whole functions.
---
### op `0x03` (`call-script`) is a raw index into the SYS4INI file table — SOLVED (2026-07-07)
The long-deferred `call-script <id>` registry (`name-resolution.md §1`) is cracked. Resolved through
the dispatch table (op `0x03``ctx[0x26c93+3]` = **`FUN_0041bc90`**), then the loader/resolver chain:
- **`FUN_0041bc90`** (handler): fetches operand 1 (the id), bounds-checks call depth (≤ 0x26), pushes
a script frame, and calls the loader.
- **`FUN_0040e980`** (loader): opens the resource by id, reads the **0x20-byte SYS4 header**, checks
magic, allocates per-frame code/local buffers from the header var-counts, reads the bytecode body,
and pushes a script frame (**stride 0x1e = 30 dwords**, indexed by `ctx[0x14f45]`). Returns to the
caller when the callee ends.
- **`FUN_0044f390`** (resolver — the key): `record = [ctx+0x414] + id*0x50`. The record is exactly the
**SYS4INI 80-byte layout** `{name[64], arc_id@0x40, file_number@0x44, offset@0x48, size@0x4c}`
(count = `[ctx+0x40c]`, archive-name table = `[ctx+0x410]`). It tries a **loose override first**
(`CreateFileA` on `record.name` → the mod/patch hook point), else opens archive
`[record.arc_id*0x100 + ctx+0x410]`, `SetFilePointer` to `record.offset`, size = `record.size`.
High-byte-tagged ids (`id & 0xff000000`) select an alternate pack via `[ctx+0x3028]` — **unused by
the corpus** (0/297 ids carry a high byte).
**So `call-script <id>` = a direct RAW index into the SYS4INI global file table** — the same table
`parse_sys4ini.py` reads, but indexed *without* skipping `@` placeholders (13208 records, 2
placeholders). There is **no separate on-disk id→code registry**; SYS4INI *is* the registry, and we
already had it. **Statically confirmed:** all **297/297** distinct corpus `call-script` ids resolve to
a `.BIN` script with a semantically-exact name (`0x1ab→ADDITEM`, `0x2ae7→MES`, `0x143→BUNKI`,
`0x329d→CALCREVISE`, `0x2add→CALCBTPARAM`), 0 out-of-range, 0 pack-branch. Tooling:
`parse_sys4ini.py` emits `build/callscript-names.json` (id→name); `sys4load` annotates
`call-script 0x1ab =ADDITEM.BIN`; the whole `build/disasm/*.asm` call graph now reads by name. See
`name-resolution.md §1`.
**Companion — op `0x8f` (`call`) is INTRA-script, not cross-script.** Its handler **`FUN_0041fba0`**
sets `[frame PC @+0x53d2c] = [frame codebase @+0x53d28] + operand*4` and pushes a return address on
the per-frame return stack (`[ctx+0x552e8]`/`[ctx+0x55248]`). The operand is a **code offset within
the current script** (matches header table **T3, tag 0x8F** = local call targets). So `0x8f` is a
local JSR; only `0x03` loads another script.
**Follow-up (functional):** the C# VM still *stubs* `call-script`. With the id→resource mapping now
known, it can be implemented for real (load the target `.BIN` from the archive via the SYS4INI record,
push a frame, run, return) — the unlock for subroutine-using scripts and, via the same path,
decision→scene (scenes are just `SCxxxx.BIN` records loaded by their SYS4INI index).
---
### op `0x215` (`query-gfx-object?`) is a native command-buffer op — settles the render drift as (b) (2026-07-07)
**This is the canonical account of the background/sprite "drift" bug** (background pinned off-centre /
bottom-right, rest grey — `Screenshot 2026-07-06 211353.png`). It supersedes the earlier "drift =
state-divergence, seed state and it's fixed" conclusion in `docs/phase-a-slice-plan.md` and the status
memory, which are corrected to point here.
Resolved via the dispatch table (`ctx[0x26c93 + 0x215]`): the registration routine `FUN_00413860` stores
`[ESI + 0x9baa0] = 0x42a0b0`, so op `0x215`'s **real handler is `FUN_0042a0b0`**. (Kelebek's `0x421160` is
VA-drift — it lands inside the unrelated `FUN_00421090`. Same lesson as `0x1a2`: never trust a Kelebek raw VA.)
`FUN_0042a0b0(ctx)` does exactly two things:
1. **`*(ctx + 0x53d88 + ctx[0x53d14]*0x78) = 5`** — writes **cmd-type 5** into the *current* gfx-object
record. A **command-buffer registration** side-effect, directly parallel to op `0x1a2` (`FUN_0042d360`)
writing cmd-type 3. So `0x215` is part of the gfx command-buffer subsystem, not a pure query.
2. **`out = FUN_0047f280(FUN_0041b940(2))`** — `FUN_0041b940(2)` fetches operand 2 (the bytecode handle
key); `FUN_0047f280` is a **`std::map::find`** over an engine-internal associative registry, returning
the mapped value or **`0xffffffff` (not-found)**; `FUN_00425fb0(1, out)` writes it to operand 1. That
registry is **populated by sibling gfx ops** — op `0x1a2`'s handler builds a `"%c%8.8x"` key and calls
`FUN_0042cf70`, an open-addressing hash **insert** into the same kind of store.
**(a) vs (b) — the verdict is (b).** The value `0x215` returns is **native command-buffer state**: "has a
gfx object already been registered under this handle?" (`≥0` = existing → use its slot; `-1` = new). That
state lives in the engine's own registry, maintained by the gfx ops, **not in the VM global bank**. So
**seeding story-state globals cannot reproduce it** — the drift is *not* the Phase-B state-divergence
problem. Stubbing `0x215` returns a constant → `label_12649`'s slot-select always takes one branch → every
draw collapses onto slot 0 → the anchor-preserve math measures foreign-sized textures → cumulative drift.
**Why the prior "state-divergence" conclusion was wrong.** It was grounded in `capture_gfx_objects.py`,
which polled the object-*record* array (`[esi+0x53d64]`) at ~2/s and saw only 3 persistent UI objects, "0
CG objects." But (i) the branch is driven by the **map lookup** (a different structure the poll never
observed), and (ii) command-buffer records are **transient** — a 2/s poll can't prove CG records weren't
used. Absence in that capture ≠ absence of the native path.
**The fix is tractable and Frida-free.** (b) does *not* mean an opaque native state machine. The subsystem
is a **modelable data structure**: an object-record array (slot / geometry / cmd-type per object) plus a
handle→object registry (a `std::map`). Geometry and draw workers lazily populate that retained-object map;
query and erase workers read/remove the same entries. Op `0x1a2` also maintains an operand-descriptor hash,
but that is a separate structure and is not what op `0x215` queries. The opcode-level source of truth is
`vm-map/opcodes.toml`.
#### Op `0x215` queries the retained gfx-object's source slot (corrected 2026-07-09)
The decisive caller/callee detail is the owner pointer. Op `0x215`'s handler passes
`ECX = ctx+0x46614` to `gfx_object_query_source_slot` (`0x47f280`); that worker searches
`ECX+0x408`. `draw-texture` passes the same owner to `gfx_object_bind_draw` (`0x47e870`), whose
`gfx_object_get_or_create` uses the same `owner+0x408` map and writes the source surface slot to
`obj+4`. Therefore:
- absent handle → `0x215` returns `0xffffffff` (-1);
- geometry-only/unbound object → its default source slot is -1;
- draw-bound object → `0x215` returns the live source slot from `obj+4`.
The earlier “op `0x215` reads a registry populated only by `0x1a2`” conclusion was wrong: it conflated
the retained-object `std::map` with `0x1a2`'s open-addressing operand-descriptor hash. The useful part of
the earlier fix remains: `GetOrCreate` must not fabricate a slot. A fresh object stays unbound (-1) until
`draw-texture` supplies its real source slot.
This also explains the reported magic-circle retention end-to-end. `AE001H.AGF` (resource `0x37`) is
bound to the ritual object's surface slot. At the post-effect cleanup (SC0000 `0x3321`), the script queries that object with
`0x215`, sign-tests the returned slot, then executes `0x1f7(handle, 10)` followed by
`0x1fa(returned_slot)`. Native `0x1f7` removes the retained object group from this same map; `0x1fa`
releases the surface slot. The port's old separate-registry model returned -1, skipped the guarded cleanup,
and left the circle compositing. `GfxState.QuerySlot` now returns `GfxObject.SourceSlot`; VM op `0x1fa`
clears that surface slot. The booted SC0000 regression ends with no visible resource `0x37`; live
clicked-path validation confirmed the corrected disappearance on 2026-07-10.
Note a **second, still-latent** gap this uncovered: `label_125bd` (which fills `rec[s3]`/`G[0x3239]` with
the eight per-object slots 4..11, called at `SC0000` `0x50f`) does **not** execute in a cold single-scene run —
the scene coroutine framework (ops `0x7b`/`0x140` + the `G[0xaba5c]==1` re-entry gate) routes cold flow
past it, so every fresh CG is assigned slot `0`. It doesn't break the *opening* (one full-screen CG shown
at a time, so sharing slot 0 is harmless and the fresh-branch geometry is correct regardless), but a scene
with several simultaneous distinct-slot objects would need the setup to run. Tracked as the scene-coroutine
work, separate from this fix.
**⇒ Scene-coroutine framework — INVESTIGATION COMPLETE (2026-07-09).** The mechanism behind the slot-0
collapse is fully understood; the native finding and the implemented host-model disposition follow:
**The gate `G[0xaba5c]` is NATIVE scene-entry state — no script sets it to 1.** Across the whole corpus
(429 references in 150 files) *every* `aba5c` reference is a read or a write of `0`; nothing anywhere writes
`1`. So `aba5c==1` is set by the engine's scene loader/scheduler on entry — the **same class as the INIT2
handle array** (native entry-state a cold single-scene harness skips), NOT a story flag. Cold, it reads `0`.
**Corrected roles of the two branches** (the earlier head-start had them inverted). At SC0000 `0x450`
`eq local0 = (aba5c==1)`; `0x457 jcc local0 label_462 <fallthrough>`:
- **`aba5c==1``label_462` "ループ開始" (loop start)** = the scene's **intro/setup LOOP**. Its body
`label_491` runs `call label_125bd` (@`0x50f`, the slot-table fill `G[0x3239..0x324e]=4..11`) plus ADV
state init, UI-slot clears (a loop over `G[0x3239]`), intro draw — then `jmp label_462` (@`0x711`). A real
loop, exited only when its iterator makes `G[0x6be]==exit-PC` (→ `mov aba5c 0`, `jmp label_45e`).
- **`aba5c!=1``label_45e → call label_71b`** = the **scene CONTENT**: `label_71b` is a `switch(G[0x62ccf])`
on the SCJUMP decision → "序章 / プロローグ", `play-bgm`, `call label_12649` (CG loads that *read* the slot
table). So the intended lifecycle is **enter `aba5c=1` → intro loop fills slots → `aba5c→0` → content uses
slots.** Cold we skip straight to content with an empty slot table → all layers collapse to slot 0 → grey.
**The loop iterator `op 0x140` is a native video-service call — not statically reproducible.** Handler =
**`0x4299c0`** (dispatch `ctx[0x9b74c]=0x4299c0`; created+typed `EngineCtx*`+annotated; Kelebek `u0041F9C0` is
VA-drift). It writes gfx cmd-type 9, copies operand-2/3 strings (`"LABEL"`, `"J"`) + operand-4 int, calls
**`(*DAT_005c6018)(8, ctx[0x54fe8], &{str,str,int})`**, and writes the returned PC-like value back to operand 1
(SC0000: `G[0x6be]`). `DAT_005c6018` is a **runtime-resolved function pointer** (all 6 xrefs are READs, no
static writer) — the engine's **native video / transition / timing service**: `FUN_00405740` (a screen-
transition/fade routine full of DirectDraw-layer calls) calls the *same* pointer with `cmd_id=3` and branches
on its return (`1`/`2` = transition progress). It is the **same class as the DirectDraw workers this project
deliberately does not model.** ⇒ faithfully emulating `0x140` = emulating the native video service = out of
static scope, permanently.
**The two companion ops (confirmed):**
- **`0x7b` (`FUN_0041ebf0`)** — *save yield handlers*: writes op1→`ctx[0x6da88+idx*4]`, op2→`ctx[0x6db28+idx*4]`
(idx=`ctx[0x53d14]`). SC0000 `0x79`: `op 0x7b label_3c9 label_41e` — registers the per-frame ADV handlers.
- **`0x7c` (`FUN_00417cb0`)** — *resume*: requires run-state bit `0x2000000` (`ctx[0x6dbc8]`) — **throws
(`__CxxThrowException`) if unset**, so it is only ever reached on a scheduler-driven re-entry, never cold;
restores PC=`ctx[0x53d28]+ctx[0x6dbcc]*4`, clears the bit, resets input/line state.
- `label_3c9`/`label_41e` (the `0x7b`-saved handlers) are per-frame *render → poll (`call label_8c`) → yield
(`0x7c`)* routines — i.e. **this coroutine machinery IS the ADV per-frame loop**, not just intro setup.
**⇒ DESIGN (host-model, not emulate).** To make cold single-scene runs correct: (1) supply `aba5c=1` as
scene-entry state (native, seedable, like INIT2); (2) give `0x140` a **host-modeled bounded "labeled yield"**
that runs the intro body at least once (so `label_125bd` fills the slot table + ADV init runs) then reports
completion so content plays — we reproduce the *observable effect*, not the video service. The idiom is
byte-identical across all ~136 ADV scenes, so it generalizes with zero per-scene work. **Payoff beyond the
slot fix:** the same `0x7b`/`0x7c` + handler machinery is the ADV frame loop, so a clean host model becomes
the seam for the interactive-ADV backlog (`0x90` hotspots / EMPTY scenes) and likely fixes the stuck magic
circle (scene-phase cleanup). **Permanently out of static scope:** the real intro-transition *timing/pacing*
inside the loop (it lives in the native video service) — we approximate it host-side, as everywhere else.
**Host model implemented (2026-07-09).** `VirtualMachine` recognizes only the ADV `"LABEL","J"` form
(138 corpus scripts; all have the same terminal-check shape), synthesizes `G[0xaba5c]=1` on top-level
scene entry, forces one setup-body pass even if `G[0x6be]` is stale, then returns the terminal immediate
discovered from the following `mov`/`eq` pair. Thus no SC0000 offset is hardcoded. Op `0x7b` retains the
saved handler PCs as frame metadata; op `0x7c` is a host-scheduler marker because `IHost.FrameYield` and
`FrameClock` already own per-frame pacing. `TITLE.BIN`'s unrelated `"BIN","SC????.BIN"` service remains
stubbed. The real video-service timing remains intentionally unmodeled.
**Magic-circle retention fixed in the host model (2026-07-09).** The effect is `AE001H.AGF`
(resource `0x37`). SC0000 already contains the correct teardown, but the port's wrong `0x215` query
returned -1 and skipped it. The corrected source-slot query now reaches `0x1f7(handle,10)` object erasure
and `0x1fa(slot)` surface release; the booted regression ends with no visible `0x37` object.
#### gfx command-buffer — op contract table (2026-07-07, full family reversed)
Every gfx op shares one shape: **write a `cmd-type` into the current object record** (`*(ctx + 0x53d88 +
ctx[0x53d14]*0x78) = <cmd>`), fetch operands via `FUN_0041b940(i)` (1-based; `docs` = the `0x1a2` variant
uses `FUN_00415f30`), then either **SET** object fields (call a native worker `FUN_0047xxxx`) or **QUERY**
object fields (write results back to output operands via `FUN_00425fb0(i, val)`). Handlers resolved through
the dispatch table (`ctx[0x26c93+op]`); all renamed in the Ghidra project `gfx_op_0x<op>_<role>`.
| op | handler | cmd | dir | argc | contract |
|---|---|---|---|---|---|
| `0x1a2` | `0x42d360` | 3 | set | 1 | operand-descriptor hash insert: key `"%c%8.8x"(3, operand-desc)``FUN_0042cf70`; separate from the retained object map |
| `0x1f7` | `0x422270` | 5 | erase | 2 | retained-object erase: `op2>1``gfx_object_erase_range(op1,op2)` erases `[op1,op1+op2)`, else `gfx_object_erase(op1)` |
| `0x1fa` | `0x4224a0` | 3 | set | 1 | release **surface slot** `ctx+0x52bd4[op1]` (vtbl free) + `FUN_00474e40(op1)` |
| `0x1ff` | `0x4227b0` | 9 | set | 4 | 3 int→float params on obj op1 → `FUN_0047e800(op1,f2,f3,f4)` |
| `0x202` | `0x4228d0` | 0xb | set | 5 | blit obj op1 with (op2,op3) + **packed ARGB** from op4(alpha)/op5(color) → `FUN_0047ea00` |
| `0x203` | `0x4229a0` | 9 | set | 4 | draw obj op1 with op2 + packed color(op3/op4) → `FUN_0047e9b0` |
| `0x212` | `0x4230c0` | 5 | set | 2 | `obj[ctx+0x14d54 + op1*4] -> +0x64 = op2` |
| `0x213` | `0x423110` | 7 | set | 3 | `obj[0x14d54+op1*4] -> +0x68 = op2 ; +0x6c = op3` (an (x,y) pair) |
| `0x215` | `0x42a0b0` | 5 | **query** | 2 | retained-object **find**(op2 handle) → op1 = obj+4 source slot / `0xffffffff`. **Drives setup and teardown.** |
| `0x216` | `0x42a0f0` | 5 | **query** | 2 | read `[ctx+0x46d14 + op2*0x14]` → op1 |
| `0x217` | `0x4231b0` | 9 | set | 4 | 3 int→float on obj op1 → `FUN_0047e960` (SETS a geom 3-vector) |
| `0x218` | `0x42a130` | 9 | **query** | 4 | `FUN_0047f360(obj op1)` → op2,op3,op4 (GETS a geom 3-vector) |
| `0x219` | `0x423240` | 9 | set | 4 | 3 int→float on obj op1 → `FUN_0047e910` (SETS a geom 3-vector) |
| `0x21a` | `0x42a1b0` | 9 | **query** | 4 | `FUN_0047f2e0(obj op1)` → op2,op3,op4 (GETS a geom 3-vector) |
**`label_12649` correlation (the drift chain, confirmed).** The recurring idiom is:
```
query-gfx-object? (G 0x62452) (G 0x6245X) ; 0x215: handle G[0x6245X] -> working slot G[0x62452]
ui-elem? (G 0x6245X) 0xa ; 0x1f7: select that element
ui-clear? (G 0x62452) ; 0x1fa: clear the slot
```
`G[0x62452]` is the **working slot**; `G[0x6245X]` are per-object **handles** (the `0x62455[idx]` family:
`0x62456/7/8/a/b/c`). The geometry ops move two per-object 3-vectors between object records and globals:
- **`0x217` SET** anchor-vector `G[0x6249b/c/d]` **into** the object; **`0x218` GET** it back **out**.
- **`0x21a` GET** position-vector into `G[0x62498/9/a]`.
These get-vectors are exactly the inputs to the anchor-preserve math (`docs/superpowers/specs/2026-07-06-a2b-graphics-geometry-design.md`:
`G[0x62498] = G[0x6249b] w/2`, foot-anchor at `G[0x6249c]`). **So the drift has two stubbed drivers, not
one:** `0x215` (wrong slot → collapse to slot 0) **and** `0x218`/`0x21a` (stale geometry vectors → the
anchor math reads garbage). Both read object state the SET ops (`0x217`/`0x219`/`0x212`/`0x213`) wrote — all
bytecode-driven, all host-modelable.
**Model implication for the host-side reimplementation (Phase 2 input).** The subsystem is a set of
per-object records keyed by handle, carrying: a live source **slot** (written by draw-texture, queried by `0x215`), a **position 3-vector**
(`0x21a` get / a matching set), and an **anchor 3-vector** (`0x218` get / `0x217` set), plus color/blit
params (`0x202`/`0x203`). The native workers (`FUN_0047xxxx` = the DirectDraw/surface layer) need **not** be
modelled — only the object-record data model, so the QUERY ops return what the SET ops stored. That makes
`0x215`/`0x216`/`0x218`/`0x21a` return correct values and the existing bytecode geometry math produces
correct `dst`/`w`/`h`. Ancillary per-object tables observed: `ctx+0x14d54` (obj pointers, fields `+0x64/
+0x68/+0x6c`), `ctx+0x46d14` (stride `0x14`), `ctx+0x52bd4` (element pointers), plus the `0x408` registry.
Worker functions decoded + annotated in the Ghidra project (updated 2026-07-09): `gfx_object_erase`(`0x47d850`),
`gfx_object_erase_range`(`0x47d8b0`), `gfx_object_query_source_slot`(`0x47f280`),
`gfx_object_get_or_create`(`0x47ddb0`, inserts a zeroed default via
`gfx_object_init_default`@`0x472810`), the setters `gfx_set_vec18/24/16c`(`0x47e960/e910/e800`), the getters
`gfx_get_vec18/24`(`0x47f360/f2e0`).
#### The `0x21c0x243` sprite transform / ANIMATION cluster (2026-07-10, partial implementation)
The scene-completeness tracker (`tools/scene_opcode_coverage.py`) flagged a dense band of GAP ops in
`0x21c0x243` (+ `0x2bd/0x2bf`) — the **largest remaining rendering unknown** in SC0000 (e.g. `0x220`×66,
`0x22f`×34, `0x228`×33, `0x21e`×25 static sites). Resolving every one through the dispatch table
(`ctx[0x26c93+op]`, read from `FUN_00413860`) shows it is **one coherent subsystem: sprite transform +
animation/tween** — and two members were already named in prior RE (`0x234 gfx_op_0x234_anim_start`,
`0x238 gfx_op_0x238_set_anim_clock`). Kelebek VAs drift here as everywhere (op `0x220` real handler is
`0x4234e0`, not Kelebek's `0x4215D0`). **Op → real handler map:**
| op | handler | op | handler | op | handler |
|---|---|---|---|---|---|
| `0x21c` | `0x417520` (417xxx trivial) | `0x229` | `0x423700` | `0x236` | `0x423ee0` |
| `0x21d` | `0x423310` | `0x22a` | `0x4237b0` | `0x237` | `0x4240a0` |
| `0x21e` | `0x423350` **✎ set_transform3_norm** | `0x22b` | `0x423850` | `0x238` | **anim_start's clock ✎** |
| `0x21f` | `0x423410` | `0x22c` | `0x423900` | `0x239` | `0x424120` |
| `0x220` | `0x4234e0` **✎ set_transform3_abs** | `0x22d` | `0x423990` | `0x23a` | `0x42a440` |
| `0x221` | `0x423590` | `0x22e` | `0x423a40` | `0x23b` | `0x424190` |
| `0x222` | `0x4235e0` | `0x22f` | `0x423b00` | `0x23c` | `0x417580` (417xxx) |
| `0x223` | `0x423620` | `0x230` | `0x423ba0` | `0x23d` | `0x4175c0` (417xxx) |
| `0x224` | `0x417550` (417xxx) | `0x231` | `0x423be0` | `0x23e` | `0x42a4a0` |
| `0x225` | `0x4236a0` | `0x232` | `0x423c30` | `0x23f` | `0x42a520` |
| `0x226` | `0x42a230` | `0x233` | `0x423cf0` | `0x240` | `0x4245f0` |
| `0x227` | `0x42a2e0` | `0x234` | **set_rotation_cycle ✎** | `0x241` | `0x4247e0` |
| `0x228` | `0x42a3a0` | `0x235` | `0x423e40` | `0x242` | `0x4249d0` |
| | | | | `0x243` | `0x4182d0` (417xxx) |
(`0x2bd``0x4251c0`, `0x2bf``0x425240`. The handful of `0x417xxx` handlers are trivial/marker-shaped — the
default-handler neighbourhood — and are almost certainly no-ops or arg-poppers; triage before modelling.)
**Contract (completed 2026-07-10, representative ops `0x21e`/`0x220`, both `argc 6`):** these are
independent matrix channels, not two encodings of one vec3 property.
- `0x21e` normalizes operands 46, then `gfx_object_set_scale_channel` (`0x47eaa0`) stores timing at
`obj+0x3c/+0x50` and calls `0x48af1d`, which writes the three values onto a 4×4 matrix diagonal at
`obj+0xac`: a **scale matrix**.
- `0x220` passes raw operands 46 to `gfx_object_set_translation_channel` (`0x47ecc0`), stores timing at
`obj+0x44/+0x58`, and calls `0x48afb1`, which writes them into matrix entries 1214 at
`obj+0x1ac`: a **translation matrix**.
- `0x21f` converts operands 47 to floats and calls `gfx_object_set_rotation_channel` (`0x47eb70`). It
stores delay/duration at `obj+0x40/+0x54`, target axis at `obj+0x1f8..0x200`, target angle (degrees)
at `obj+0x208`, and the target axis-angle matrix at `obj+0x12c`. Current axis/angle are
`obj+0x1ec..0x1f4/+0x204`, with current matrix `obj+0xec`.
- `gfx_object_apply_transform_channels` (`0x472f00`) supplies the timing contract. All three channels use
shared start timestamp `obj+0x34` and global frame-time `ctx+0xb550`, but have independent delay/duration:
scale `obj+0x3c/+0x50`, rotation `obj+0x40/+0x54`, translation `obj+0x44/+0x58`. Each holds current through the
delay, linearly interpolates current→target for its duration, then commits the target and clears its own timing.
Neither third component is opacity.
**Exact composition and 2D reduction (live-validated 2026-07-10).** The one-shot consumer starts from identity and
right-multiplies `T(-V18) → scale-current → rotation-current → translation-current → T(+V18)`;
`matrix4_multiply` at `0x4ee2a4` computes `out = left * right`. AGE uses row vectors. With no
rotation/perspective, the screen projection is therefore exactly
`V18.xy + (point.xy - V18.xy) * scale.xy + translation.xy`. The captured SC0000 handle `0xcbc0`
has base `(0,600)`, anchor `(400,1000)`, and final scale `(5,5)`; native matrix translation
terms are `(-1600,-4000)`, projecting the base point to `(-1600,-1000)`. The port's focused
projection test and transform-aware gfx log reproduce those values.
`gfx_object_composite` then right-multiplies `gfx_object_anim_interpolate`'s separately anchored product,
which contains op `0x234`'s cyclic rotation. With the other oscillating matrices at identity, adjacent anchors
cancel and the full order is
`T(-V18) * scale * one-shot-rotation * translation * cyclic-rotation * T(+V18)`. Thus cyclic rotation
also rotates the translation vector. The cyclic angle is integer degrees
`floor(((frameTime-start) % period) * 360 / period)`; it wraps to zero without ping-pong. Positive Z produces
`m01=+sin, m10=-sin`, clockwise on the Y-down screen.
Native matrix oracle: handle `0xcb8e`, anchor `(700,600)`, scale current `0.9`, op `0x21f` target axis
`(0,0,1)`/30° after 500 ms for 390 ms, sampled 11 ms into the ramp as
`[0.9055,0.0134;-0.0134,0.9055]` with translation `(74.1449,47.3127)`. The port focused test matches
those terms. In the windowed port capture, the two SC0000 `0x234` sites (periods 9000/13000 ms, Z axes
`+1/-1`) advanced after 563 ms to integer angles `22/15`, exactly the native formula, and produced distinct
affine PNG frames. Nearest-neighbour inverse mapping is the deliberate software raster sampling policy;
native D3D9 subpixel filtering remains a possible pixel-level difference, not an uncertain matrix approximation.
**Port result (2026-07-10):** `GfxState` retains scale, one-shot rotation, translation, and cyclic rotation
with their native clocks/order. `Transform2DMath` composes the full row-vector 4×4 transform before 2D
projection. Godot uses an inverse-mapped affine RGBA8 rasterizer for textured objects and solid fills,
preserving colorkey/tint/opacity behavior and never deriving opacity from transform Z.
##### `anim_start`/`set_anim_clock` decoded + opening confirmed (2026-07-07, animation-slice Task 1)
Decoding the two already-named clock/start ops (dispatch table → `0x234`@`0x00423da0`, `0x238`@`0x004240e0`;
both annotated) and grepping the SC0000 opening settles the animation model and confirms the opening exercises it:
- **`0x238 set_anim_clock` (argc 1, cmd-type 3):** `ctx+0x51b78 = 0` (elapsed), `ctx+0x51b7c = operand1`
(total duration). **A GLOBAL, NON-BLOCKING clock** — not per-object. The op only *configures* the clock; it
does **not** loop/wait. The native render loop advances this clock each frame and interpolates *all* animating
objects. Its own plate comment states the payoff: "our port can drive animation in the host's per-frame loop
while the VM is parked at wait-for-input; no VM/host frame-lockstep." → **validates the wall-clock-tween
architecture directly.** SC0000: `set-anim-clock(G[0x624bb])` @`0x123bd`, `set-anim-clock(0x190=400)` @`0x13858`.
- **`0x234 anim_start` (legacy mnemonic; argc 5, cmd-type 0xb):** following worker
`gfx_object_set_rotation_cycle` (`0x47f060`) into `gfx_object_anim_interpolate` (`0x473ed0`)
corrects its ABI to `(handle)(period_ms)(axis_x)(axis_y)(axis_z)`. Period is `obj+0x228`, axis is
`obj+0x244..0x24c`, and the frame-clock consumer applies
`360*((now-start)%period)/period` degrees. This is a cyclic **rotation** channel, not a target vec3
for scale/translation and not opacity.
- **The opening path uses the whole subsystem, early.** `0x21e`/`0x220` transform-sets fire from `0x00f73`
onward (`0x21e (G[0x6245b]) 0 0x12c l0 l1 0x64`, `0x220 (G[0x62457]) 0x96 0x3e8 l1 l3 0`), on the same INIT2 CG
handles (`G[0x62457]`,`0x6245b`,`0x6245c`) — this is the opening, **not** battle/debug. So the slice's ops are
real and verifiable on screen.
**Corrected host model:** `0x21e` scale and `0x220` translation run directly from frame-time
`ctx+0xb550` using their own delay/duration; they do not use op `0x238` as their duration.
Op `0x234` is the independent rotation cycle above. Op `0x238` still configures the separate
`ctx+0x51b78/+0x51b7c` animation service used by its own family.
##### `label_1235a` animation-section boundary (2026-07-10)
The section helper computes the maximum configured duration and arms it with `0x238`, then reads
message-skip through `0x1c7` and ADV read/click-skip state through `0x1cc`. The raw `jcc` order matters:
normal playback (both zero) branches to `0x21c` (set run-state bit `0x400`), while a nonzero skip/read
state executes `0x243` (reset the separate global animation-service clock) followed by `0x20c` present.
Both branches finish with
`0x224`, which clears the native gfx command queue at `ctx+0x418`. These handlers are now named,
commented, and saved in the Ghidra image.
The matching native presentation trace corrects the earlier cadence model. Ordinary opcode execution is
**burst-fast between presentation services**, while `0x21c` is the render/wait boundary: it parks the
interpreter and `gfx_render_frame` repeatedly samples visible finite one-shot channels and queued surface
commands until dirty presentation state clears; `0x224` then clears the command queue. `0x20c` is a single
explicit publication on the skip branch.
At the opening AE001D passage, native `0x125a6` rendered the preceding state, then both object binds plus
mode-1 `0x203` and target `0x202` writes (`0xd5a/0xd63/0xd73/0xd8a`) completed in about 5 ms with **no render
between them**. Their first composition was the following repeated `gfx_render_frame` loop at `0x21c`.
Likewise, the explicitly presented mode-0 white CG at `0x125a6` survived only about 10 ms before that next
boundary. The port's former 200-completed-op/s throttle stretched the same burst across many display frames;
that average had folded service waits into execution time and was not an opcode scheduler rate.
The Godot host therefore leaves ordinary `FrameYield` non-blocking and publishes retained mutations only at
`0x20c`, `0x21c`, sleep, and stable input waits. `0x21c` waits for visible finite color/matrix channels plus
`0x223`; ambient cyclic/spritesheet pulses do not block, and click forcing remains restricted to `0x223`.
This is a native-evidenced scheduler correction, not a guessed duration sleep.
##### The opening render path is RETAINED, not immediate-mode (2026-07-08, ground-truth correction)
A working note in the animation slice mis-called the SC0000 opening a set of "immediate-mode slot-0 blits." That
was **wrong**, and it came from trusting our own `Age.Cli gfx` oracle (which executes our VM and mis-labeled the
CG draws as "slot 0"). Verified against native code + the raw bytecode:
- **`draw-texture` (op `0x1fb`, handler `gfx_op_0x1fb_draw_bind`@`0x422510`) is a RETAINED bind, not a blit.** It
writes cmd-type `0x11` and calls **`gfx_object_bind_draw`@`0x47e870`**, which on the object keyed by `handle`
(operand 1) sets: `flag|=1` (visible), `obj+4 = source SLOT index`, `obj+8..0x14 = source rect`,
`obj+0x24/28/2c = position`. Its plate comment (prior RE) already states the key fact: the object stores the
**slot INDEX — a live ref to `surface[slot]`, resolved each frame at render — NOT a texture snapshot.** Objects
persist and are composited each frame; this is exactly the surfaces+objects model in "The full gfx render
model" above.
- **The SC0000 opening is a retained scene of distinct objects, `sleep`-paced.** Raw bytecode: fixed-handle UI
objects (`0xcf08` slot 3 full-screen, `0xc350` slot 0xe, `0xe678` slot 0xd — a 400×30 element re-bound 20+
times), an animated sprite (`draw-texture (G[0x62457]) (G[0x62452]) … (G[0x62498]) (G[0x62499])`, computed
position), and the CG loader (`SC0000` @`0x126e1`/`0x12970`): `set-texture G[0x62424] → slot G[0x62452]`,
`get-texture-size`, centre it, then `draw-texture (handle = CG_array[G[0x62450]] = INIT2 array G[0x62455..])
slot G[0x62452] …`. `sleep 0x64/0x3e8/0x2ee` sits between steps. So different draws use **different handles and
per-object working slots** — not one slot-0 canvas.
- **Why our port still doesn't animate the opening (conclusion unchanged, mechanism corrected):** we execute the
whole load/draw/`sleep` sequence **instantly** — no `sleep` timing, no per-frame present — so we only ever see
the *final* retained state; the intermediate `AE*` frames (`AE001D→AE002B→AE003B`, surface swaps on the working
slot between paced frames) never get a frame to display. The fix is **frame-pacing** (scene-coroutine / `sleep
0xc8`), a separate subsystem from the transform/alpha channel. **Lesson: never characterise the engine's render
mechanism from our own VM's oracle output — use native code + raw bytecode.**
### `sleep` (op `0xc8`) — the frame-pacing primitive (2026-07-08, decoded)
Handler resolved via the dispatch table (`ctx[0x26c93+0xc8]` = `param_1[0x26d5b]` in `FUN_00413860`) →
**`sleep_op_0xc8`@`0x420ec0`** (was `LAB_00420ec0`; created + annotated). It is **NON-BLOCKING**:
- It **arms a timer**`sleep_timer_arm`@`0x44cff0` on the object at `ctx+0x5f304`: `+8 = 1` (active),
`+0x14 = (*DAT_0056f3d4)()` (start tick — an **ms** source, `timeGetTime`/`GetTickCount` class, same
`DAT_0056f3d4` the boot uses to seed `srand` via `time/100`), `+0x18 = duration` (operand, min 1). The engine's
main loop polls `elapsed ≥ duration` and resumes the script — rendering continues in the meantime. This is the
native confirmation that the engine paces animation in its per-frame loop, not by blocking.
- **Operand unit = MILLISECONDS.** `duration < 10` fast-paths through import `[0x56f0b8]`; every real scene sleep
(`100`/`750`/`1000` in SC0000) is `≥ 10` → the timer-arm path.
- The handler also writes gfx **cmd-type 3** into the current object record (`ctx+0x53d88+curidx*0x78`) and runs
two **anti-tamper** checks (call `[ctx+0x5512c]`; a rotate-checksum compare of `ctx+0x55120/0x55124`;
`__CxxThrowException` on mismatch — integrity work piggybacked on a hot op). Neither is needed by our model.
**Port equivalent (implemented):** our VM runs on a background thread (like `wait-for-input`), so blocking that
thread for `duration` ms while the main-thread compositor (`Main.Recomposite` in `_Process`) keeps presenting is
behaviorally equivalent to the native non-blocking timer. This correctly reproduces the **explicit one-shot
sleeps** (the dramatic 1000/750/200 ms holds). Headless/CLI hosts no-op `Sleep` (parity). `IHost.Sleep(long)` +
VM `case "sleep"`; see `vm-map/opcodes.toml` 0xc8. **⚠ It does NOT make the rapid opening `AE*` burst animate**
— execution trace shows the back-to-back `set-texture→draw-texture` swaps have no `sleep`/`wait`/`present`/
coroutine between them, so bare `sleep` was never their pacer; what advances that burst is still unknown (an
earlier claim that "the opening is sleep-paced" was inherited from this doc and never execution-verified —
corrected). Profile the real Godot run (`--trace-histogram`) to find it.
**Related — `present-frame` (op `0x20c`):** dispatch `param_1[0x26e9f] = gfx_op_0x20c_present_frame`
`gfx_render_frame`@`0x4820b0` (buffer flip). Our compositor presents every frame regardless, so `0x20c` is a VM
no-op (`noop_headless=true`); the Kelebek label `u00416200` was VA-drift. This corrects the earlier open item
("no per-frame present") above — present is host-implicit; only `sleep` timing was missing.
### Frame cadence — the interpreter tick, and why our port "speeds through" (2026-07-08)
Answers the open question the `sleep` section above left ("what advances the rapid opening burst is still
unknown"). The pace is an **engine-level execution cadence**, not any bytecode primitive. Corroborated in-game
by Ctrl fast-forwarding ADV (a speed governor). Motivated by the user's observation that our port visibly
speeds through the opening — which contradicted, and correctly overturned, an earlier same-day overclaim that
"there is no missing pacer" (that was inferred from headless op-counts, which cannot render).
**Confirmed from the engine image (annotated in Ghidra):**
- **The interpreter is a cooperative one-op-per-tick step, not a run-to-completion loop.**
`adv_interpreter_tick`@`0x410fb0` (renamed from `FUN_00410fb0`) executes **exactly one opcode** per call:
`op = **(ctx+0x53d2c + curCtx*0x78)`; if `0 ≤ op ≤ 0x3ff` it dispatches `(*(ctx+0x9b24c+op*4))()` (the
handler table = `ctx[0x26c93+op]`) then advances `PC += *(ctx+0x53d88+curCtx*0x78) * 4` (decoded cmd size),
else the default handler `FUN_004162b0`. It also runs the **message-skip / click / auto-advance** logic each
tick (`s_set_CancelMesSkipOnClick`, `s_message_ReadTextSkip`, skip bit `ctx+0xa0ce4 & 0x8000000`) — i.e. the
**Ctrl fast-forward governor lives at the per-op level**, and a click can reposition the PC (skip-to-next).
- **Script contexts are coroutine records.** `curCtx = *(ctx+0x53d14)` indexes `0x78`-byte records at
`ctx+0x53d60`/`ctx+0x53d2c` (PC, codebase, cmd-size). The engine multiplexes script "threads." Init/reset =
`scene_context_init_reset`@`0x40b3b0` (zeroes `0x53d14` + `0xa0ce4`, allocs surfaces `ctx+0x52bd4[1000]`).
- **Advancement is gated by an interpreter run-state flags word `ctx+0xa0ce4`** (bit1 = sleeping, plus wait/
skip/etc.), read+written by ~40 state functions. `sleep_op_0xc8` sets bit1 + arms the ms timer and returns —
it does not block. So the outer loop consults `0xa0ce4` to decide whether to step the script this frame.
- **Effects are frame-stepped.** Screen transitions `FUN_0043cdb0` (12 wipe/slide modes) render **one frame per
step** and take a step-count parameter (the natural place a speed multiplier applies); `present-frame` (0x20c)
and the anim clock (0x238) advance per frame. A CG transition therefore spreads over many real frames.
- **Timing source** = the ms-clock function pointer `*DAT_0056f3d4` (`timeGetTime`-class), used throughout.
**Model:** single-threaded, vsync-timed frame loop; each frame it steps opcodes until the context **yields**
(`sleep` armed / `wait-for-input` 0x72 / active frame-stepped transition/anim / present), renders
(`gfx_render_frame`), waits on the clock, continues. Back-to-back draws inside one page compose into a single
frame (fine); the opening's CG-to-CG advances are gated by frame-stepped transitions + sleeps, which spread
them over real time.
**⚠ Not statically resolvable (honest boundary):** the **outer frame loop itself** is not readable from this
dump. `adv_interpreter_tick` is invoked through a **runtime-set mode function pointer** (heap/vtable slot) — it
has zero static xrefs, and its address bytes (`b0 10 41 00`) appear nowhere in `range_00400000` (0x400000
0x65ffff). The functions touching the scheduler state (`0x53d14`, `0xa0ce4`) are init/reset, save
(`context_state_serialize`@`0x40d320`), and op-handlers — never the loop. The "run-until-yield then render"
statement above is a **reconstruction** from those pieces, not a line read from the loop; pinning the actual
loop + its exact per-frame step budget / vsync wait needs a **live-debugger break** (attach + break in the
frame loop), or a wider memory dump that includes the mode object.
**Port relevance (the speed-through root cause).** Our Godot VM runs on a **free-running background thread**
(`Task.Run(() => vm.Run())` in `Main.cs`) with no frame binding — it executes an entire page's ops in
microseconds; only `WaitForInput` and `Sleep` pause it, and the compositor merely samples `GfxState` at 60fps.
So every no-`sleep` CG/state advance collapses to its end state → the speed-through. **Fix shape:** throttle
the VM to a bounded wall-clock op rate (see the measured numbers below) via a per-opcode host yield; retire the
free-running thread. Spec: `docs/superpowers/specs/2026-07-08-frame-stepped-vm-design.md`.
### Frame cadence — live measurement (2026-07-08, Frida read-only)
The static pass couldn't reach the outer loop, so we measured the running game. **Read-only / import-only
only** (`tools/frida/probe_frame_cadence.py`, `probe_present.py`): a plain-JS hook on the proven operand-fetch
`0x41b940` (grab ctx + count exec rate) + system-DLL hooks; no engine-code patching. **Lesson learned the hard
way:** a first attempt with a **CModule** hook on the hyper-hot `adv_interpreter_tick` crashed the game
instantly (bad native callback into the hottest path — *not* anti-tamper; our other scripts hook engine code
via plain JS and survive). Use plain-JS hooks on proven addresses + memory polling.
Findings:
- **Execution is rate-limited, not free-running.** Normal active rate ≈ **1,788 operand-fetches/sec** (peak
~5,796) — far below an unthrottled interpreter (millions/sec), so the engine paces itself. Execution is
bursty (parked at `wait-for-input` prompts, then a bounded burst), confirming per-iteration op-budgeting.
- **Fast-forward (Ctrl) scales the rate ~4×** (≈7,738/sec avg, peak ~15,572), gated by the engine skip bit
**`ctx+0xa0ce4 & 0x8000000`** (set only while fast-forwarding). It runs *more ops per unit time* — it does
not skip content. (Ctrl is **ADV-scoped**; it does not speed up gameplay/menus.)
- **Rendering = Direct3D 9, UNCAPPED.** `ddraw.dll` is not loaded; the game uses `d3d9.dll` (+ `nvd3dum.dll`).
`IDirect3DDevice9::Present` (device vtable slot 17, found by scanning ctx for a d3d9-vtable object with a
full ~119-method table) fires ~**1,908/sec** with **no vsync**; `BeginScene`/`EndScene` never fire → a **2D
StretchRect-style compositor**, not a 3D scene. So there is **no fixed display-frame rate**; `Present` rate
≈ op rate (~1 op per present). ⇒ the pacing quantity is the **wall-clock op rate**, not a per-frame budget.
- **Corrected implication for the port (2026-07-10):** 1,788/s is an operand-fetch rate, not an opcode rate.
`IHost.FrameYield` runs once per completed opcode, so matching those quantities directly overclocked the
port by roughly ninefold. The native SC0000 transform lifetime pins the corresponding completed-op cadence
at about **200/s**. The limiter must be wall-clock based (not a per-render callback budget), with a future
ADV Ctrl multiplier applied to the same unified clock.
### The render drift's SECOND half: missing system-boot state (2026-07-07, resolved)
Implementing the gfx ops (above) was necessary but not sufficient — a cold single-scene run of SC0000 still
drifted. Runtime tracing found **the CG handle array `G[0x62455..0x6245c]` was all zeros**, so every CG
collapsed onto object `0` and its geometry accumulated. Those handles are set by the **boot script `INIT2`**
(mov `0x62455=0xcb20 … 0x6245c=0xcbc0`), which is call-scripted by the real entrypoint **`SYSTEM4.BIN`**
(`LOADCONFIG → INITCONFIG → INIT2 → LOGO → OP → INIT → TITLE → …`). Our harness teleports straight into
SC0000, skipping that boot. **Fix:** run the system-boot state prefix (`INITCONFIG/INIT2/INIT`, skipping the
UI scripts) before the scene — `Age.Cli gfx --boot` and Godot `--boot` (both via `GameSession`). With boot,
the CGs **de-collapse and render correctly** (screenshot-confirmed). **This is the synthesis of the old
(a)-vs-(b) debate: the drift needed BOTH the native gfx ops (b) AND boot state (a) — specifically INIT2's
handle array, never before identified (it is not a story flag).** Note two distinct boots: our Phase-B
`--boot` runs the *data* `*INIT` scripts (skills/items/…); this is the *system* boot (`SYSTEM4` prefix) — a
"full boot" should run both. **Residual:** the `AE*` fade/flash effects still draw opaque (alpha/blend
deferred — Phase 2 scope), and some object-slot CGs start with a zero anchor (cold gfx objects vs the real
game's warm ones; default object geometry is confirmed `(0,0)` in `gfx_object_init_default`, so it is not a
missing-default bug). See `docs/phase-a-slice-plan.md` A2b-Geometry.
### The gfx animation/effects subsystem — the `AE*` fades (2026-07-07)
The `AE*` flash/glow effects (and sprite motion) are a **native time-animated retained render loop**, not
per-frame bytecode. Reversed + annotated in Ghidra:
- **Retained objects carry cyclic rotation state:** flag value `4` at `obj+0`, start timestamp
`obj+0x214`, period `obj+0x228`, and axis `obj+0x244/248/24c`.
- **`gfx_object_set_rotation_cycle`** (`0x47f060`, worker for legacy op **`0x234`**)
configures that channel. `gfx_object_anim_interpolate` consumes it from frame-time `ctx+0xb550`
as a repeating 0..360-degree axis rotation. Op **`0x1fd`** is a separate scaled 3-vector setter.
- **Op `0x238`** (`gfx_op_0x238_set_anim_clock`) sets a **global animation clock**, **non-blocking**:
`ctx+0x51b78 = 0` (elapsed), `ctx+0x51b7c = duration` (the max per-object duration; SC0000 `label_1235a`
maxes a table to compute it). It does **not** loop/wait.
- **Frame model:** the bytecode does `configure anims (0x234/0x1fd) → set clock (0x238) → show-text →
wait-for-input` and **continues**; the native render loop advances the clock + per-object progress each
frame, interpolates, composites, presents. Render/present family nearby: `0x243/0x20c/0x21c/0x224`
(`u004162xx`, not yet fully RE'd). **⇒ the port can drive animation in the HOST per-frame loop while the
VM is parked at wait-for-input — no blocking present op, no VM/host frame-lockstep** (the answer to the
"frame loop" question).
Consequence: animation needs a retained per-frame compositor. That architecture is live; scale,
one-shot rotation/translation, and cyclic rotation now rasterize through the affine software path.
### The full gfx render model — surfaces + objects + composite (2026-07-07)
Reversed the create/set/draw-texture handlers + the render loop (all annotated in Ghidra). **This is the
canonical model** (an earlier flat "draw layers to one screen" attempt was WRONG — it had no surface concept
and snapshotted textures at draw time; symptoms: alternating grey, glow over backgrounds, vanishing sprites).
**Two distinct stores:**
- **Surfaces** — image buffers at `ctx+0x52bd4[slot]`, indexed by slot. `gfx_op_0x1f8_create_surface`
(`0x4222d0`) allocates a blank one (releasing any old); `gfx_op_0x1f9_load_surface` (`0x422360`, op `0x1f9`
set-texture) resolves `resId` via the SYS4INI resolver (`FUN_0044f390`) and loads the file into the slot's
surface **with a colorkey/chromakey** (op arg 3 — never modelled before), also releasing the old surface.
A surface persists at its slot until the next set-texture overwrites it.
- **Objects** — the `ctx+0x408` registry, keyed by handle (a `std::map`). `gfx_op_0x1fb_draw_bind` (`0x422510`,
op `0x1fb` draw-texture) → `gfx_object_bind_draw` (`0x47e870`): sets the object's **source slot** (`obj+4`),
**source rect** (`obj+8..0x14` = left,top,right,bottom), **position** (`obj+0x24/28/2c` = V24), and the
**visible** flag (bit 0). The object references its surface **by slot index, live** (re-resolved each frame),
NOT a snapshot. Objects also carry anchor V18 (`obj+0x18`), independent scale and translation
matrices/timing, cyclic rotation state, and color/alpha (`0x202/0x203`).
**Render frame** — `gfx_render_frame` (`0x4820b0`), driven by op `0x20c` present (`gfx_op_0x20c_present_frame`
`0x4174a0`, which also updates the frame timer `ctx+0x51b64/68`): iterate the object registry **in ascending
handle order — that IS the z-order** (lower handle behind, higher on top; `std::map` key order). For each
object with visible bit 0, `gfx_object_composite` (`0x47f650`) computes its transform from geometry, **applies
the animation interpolation if bit 2 is set**, and blits `surface[obj.slot]` with alpha/colorkey. Then swap
buffers (present). **Slot 0 is NOT special** — a normal slot; several objects may share one surface.
**⇒ Faithful port:** a `SurfaceStore` (`slot → {image, colorkey}`, from create/set-texture) + an `ObjectStore`
(`handle → {slot, srcRect, position, anchor, scale, anim, alpha, visible}`, from draw-texture + the gfx ops) +
a host per-frame compositor that draws visible objects **in ascending-handle order** from their live surface,
interpolating animations by elapsed time. No VM/host lockstep: animations play during the wait-for-input park.
Separate scale/rotation/translation state and timing are implemented. Anchor semantics, multiplication order,
cyclic wrapping, 2D projection, and affine raster coverage have focused native-oracle tests. Native D3D9 filtering
and render-target command execution remain separate fidelity work.
### Blend & transparency — colorkey + `0x202`/`0x203` color/alpha (2026-07-08)
Reversed for graphics slice A (spec `docs/superpowers/specs/2026-07-08-blend-transparency-design.md`;
Ghidra functions renamed + plate-commented, saved).
- **Colorkey format** (`gfx_op_0x1f9_load_surface` `0x422360`): read op arg 3; if `(int)key < 0` → **no
colorkey** (opaque); else the operand is **`0xRRGGBB`**, converted to `0xFFRRGGBB` and passed to the
surface creator `FUN_00477c40` as the transparent key (so operand `0` = key **black**). Colorkey is baked
**at surface-load** (matching texels → transparent), NOT compared per-blit. *Port:* interpret arg 3 as
RGB888; `<0` = none; else texels whose `(R,G,B)` equal the key become transparent when the surface image
is loaded/cached.
- **`0x202` (`gfx_op_0x202_worker_set_color_anim` `0x47ea00`)**: sets an **animated** color/alpha target
`obj+0x64 = packedARGB`, the color-anim active bit, and resets shared start `obj+0x34=0`. Operands 2/3
are delay/duration at `obj+0x38/+0x4c`; sampling uses frame clock `ctx+0xb550`, not op `0x238`.
- **`0x203` (`gfx_op_0x203_worker_set_color` `0x47e9b0`)**: sets a **static** color/alpha `obj+0x60`, no anim
bit. Immediate per-object modulation.
- **Blit** (`gfx_object_blit_d3d9` `0x4774c0`): selects a **blend mode** (`local_2c`: 0 opaque, 1 alpha
`SRCALPHA/INVSRCALPHA`, 2/3 additive/special for glow/flash) and passes a modulation color/alpha to the
device draw. Slice A ports the **alpha** path (fades); additive (glow) is deferred (its `local_2c` source
field is `obj+0x30`, the value written by op `0x203`. Mode 1 uses packed ARGB alpha as opacity and RGB
as multiplicative D3D modulation. Modes 2/3 remain separately scoped beyond the completed mode-1 path.
**Interpolation RE pass (2026-07-08, stalled → both deferrals confirmed).** Attempted to pin how a `0x202`
fade animates so smooth ramping could join slice A. Findings (Ghidra `gfx_object_anim_interpolate`
`0x473ed0`, annotated+saved): the bit-2 anim family (op `0x234`) interpolates **5 independent sub-channels**
(color `obj+0x240`/period `obj+0x220`, two matrices, rotation `0x168`=360°, src-rect scroll), each on the
**global frame clock `ctx+0xb550`** (advanced per present, NOT the op-`0x238` clock `ctx+0x51b7c`), and each
**ping-pongs** (triangle wave, folded at period/2) — i.e. these are *oscillating/pulsing* effects, not
one-shot fades. The `0x202`/`0x203` color (`obj+0x60` static / `obj+0x64` animated, sets
`ctx+0xb558`/`0xb560`) is a **separate channel whose blit consumer was not located** in this pass. So a
one-shot fade's exact source→target→easing is still unresolved and would take a dedicated dig (find the
`obj+0x64` consumer + the color→`obj+0x240` path + the clock advance). ⇒ **smooth color-anim interpolation
stays deferred**; slice A ships the static end-state (which reaches the correct final alpha/tint and fixes
the stuck-opaque bug), with interpolation as a scoped follow-up.
**Resolution (2026-07-10 — supersedes the deferral above).** The missing consumer was the bit-1 branch in
`gfx_object_apply_transform_channels` (`0x472f00`), before its matrix channels. It seeds shared start
`obj+0x34` from `ctx+0xb550`; holds current packed ARGB `obj+0x60` through delay `+0x38`; then performs an
integer, bytewise linear interpolation to target `+0x64` for duration `+0x4c`. At natural completion—or
when `ctx+0xb55c == 1` requests forced completion—the target commits to current, delay/duration clear,
target becomes `0xffffffff`, and the one-shot active bit clears when no color/matrix/src-rect sibling remains.
The object-local override bit at `+0x2d0` suppresses the global force. Negative alpha/RGB target operands
independently preserve their bytes from current `+0x60`.
The port now carries current and target separately and samples them from the unified `FrameClock`; an op
`0x203` static write after `0x202` therefore becomes the ramp's current value rather than overwriting its
target. Mode 0 retains the established CG/tint/fill behavior; mode 1 now uses native alpha opacity plus RGB
modulation. `draw-string 0x204`/`0x7a` remains a separate dependency.
### SC0000 anim/transform/spritesheet cluster — op→field map (2026-07-08)
Reversed for the animation cluster slice (spec `docs/superpowers/specs/2026-07-08-sc0000-anim-transform-cluster-design.md`).
Every cluster handler resolved via the dispatch table `handler(op)=ctx[0x26c93+op]` (the `opcodes.toml`
`u004xxxx` labels are Kelebek VA drift — do not use them). Each op is a thin wrapper (`FUN_0041b940(n)`
fetches operand n) → a worker that writes object fields; the interpolator `gfx_object_anim_interpolate`
(`0x473ed0`) is the consumer. **The cluster is heterogeneous** — setters, queries, and a movie op. Renamed +
annotated in Ghidra, saved.
**In scope (built this slice):**
| op | handler / worker | semantics |
|---|---|---|
| `0x22f` | `gfx_op_0x22f_set_position_anim` → `gfx_worker_set_translation` | set object **position** (translation vec `obj+0x5d4`); base transform, not a ping-pong channel |
| `0x229` | `gfx_op_0x229_set_position` (`FUN_00472bb0`+`FUN_00472be0`) | set object **position/geometry** immediately (`obj+0x420/0x424` + vec `obj+0x440..0x448`) |
| `0x239` | `gfx_op_0x239_set_srcrect_cell` → `gfx_worker_set_srcrect_cell` | set **spritesheet** grid `obj+0x238/0x23c` + static cell `obj+0x234` |
| `0x231` | `gfx_op_0x231_anim_srcrect` → `gfx_worker_anim_srcrect` | **animate spritesheet cell**: bit2 active, period `obj+0x230`, grid `obj+0x238/0x23c` → interpolator SRC-RECT SCROLL channel (ping-pong across the row) |
| `0x232` | `gfx_op_0x232_anim_color` → `gfx_worker_anim_color` | **animate color** (pulsing GLOW): bit2 active, period `obj+0x220`, target `obj+0x240` → interpolator COLOR channel (ping-pong). Distinct from static `0x202`/`0x203` (`obj+0x60/0x64`) |
| `0x228` | `gfx_op_0x228_query_position` (`FUN_0047cdd0`) | **query** current computed (x,y,z) → operand slots 3/4/5 (script logic, not render) |
| `0x23f` | `gfx_op_0x23f_query_object` (`FUN_0042a520`) | **query** an object status/value → operand slot 1 |
**Follow-up resolution (2026-07-10):** `0x21f` is the one-shot axis-angle channel and is implemented with
affine rasterization. `0x223` is **not affine**: `gfx_queue_surface_alpha_transition` (`0x47f440`) inserts
a type-0 command-map record keyed by arg 1: start `+4`, delay/duration `+8/+0xc`, target surface slot `+0x10`,
and two object handle ranges at `+0x14/+0x1c` and `+0x18/+0x20`. `gfx_render_frame` composites those ranges
into the target and ramps alpha 0→1. Its SC0000 site `0x129e7` passes `(handle+2, transition slot,
handle+1,1,handle,1,G[0x6249f],G[0x624a0])`. It remains a render-target/transition slice dependency rather
than being approximated in the affine object compositor.
#### ADV foreground surface-transition lifecycle (2026-07-10)
The completed native chain is `0x21d` snapshot -> `0x223` queue -> `label_1235a` skip queries ->
`0x21c` normal yield or `0x20c` skip endpoint-present:
- `gfx_object_clone` (`0x47e4f0`, op `0x21d` handler `0x423310`) copies the complete retained-object
record: exactly `0xb5` dwords / `0x2d4` bytes. SC0000 `0x128fc` clones the current CG handle to
`handle+1`; the loader then rebinds the source handle to the new CG, and `0x223 @ 0x129e7` uses
`handle+1` as old range A, the updated handle as new range B, and `handle+2` as the target presenter.
- `0x1c7` is the `run_state_flags & 0x08000000` message-skip query. `0x1cc` reads
`ctx+0x6dbd4`, now named `adv_read_skip_state`; `adv_refresh_read_skip_state` (`0x406cd0`) and the
text/label/wait handlers maintain it from `message_ReadTextSkip` plus per-PC read history. It is not
surface-transition progress.
- Zero OR-state is normal playback and reaches `0x21c`; run-state bit `0x400` parks interpreter
progression while the type-0 surface command advances from the frame clock. Nonzero skip/read state
reaches `0x243 + 0x20c`, exposing the completed endpoint without the normal wait.
- Op `0x203` stores operand 2 at object `+0x30`. Transition sources use mode 2; their common
`0xffffffff` color is opaque identity modulation, not a solid-white tint. Negative alpha/RGB operands
separately preserve the corresponding byte(s) of current static color via the `FUN_0047f3e0` lookup.
The port mirrors this with an explicit pending/start/progress/natural-or-forced-complete state. A click
while `0x21c` is parked completes only the active foreground surface transition and is consumed; it does
not pre-arm the following `wait-for-input`, and it does not complete independent retained rotation,
matrix, spritesheet, or color-animation channels.
**Still deferred:** `0x236` (`gfx_op_0x236` @`0x423ee0`) a
**timed/animated-surface (movie-like) op**; plus the unclassified `0x242/0x23d/0x20a/0x20e`
tail (2-arg flags / inline). These stay GAP until a follow-up slice or are safe-noop'd if the opening tolerates it.
### Grey-background root cause — slot collision + tint-strength (2026-07-08, gfx-log)
Diagnosed with the new `--gfx-log` compositor/op trace (docs/tools-reference.md). The grey background has
**two distinct causes**, both now proven:
1. **CORRECTION to the blend section above — op `0x202`/`0x203` "alpha" is a TINT STRENGTH, not object
opacity.** Evidence: the primary CG is drawn with `0x203 (alpha=0, color=white)` = `0x00ffffff`. That
means "blend the tint (white) into the texel by strength 0" = **no tint, fully opaque CG** — but slice-A
treated the alpha byte as the object's opacity → the CG rendered fully transparent → grey. Fix (commit
5e4fdda): `RenderObject.TintStrength` split from `Alpha` (opacity); textured objects stay opaque and the
tint LERPs the RGB by strength (0=keep texel, 1=full tint). Surfaceless fills use the strength as fill
opacity. Verified: the opening event CGs render again (shot-confirmed).
2. **Effect pages: everything collapses into slot 0.** `set-texture` is dominantly
`set-texture (GLOBAL resId)(GLOBAL slot)(local colorkey)` (543× across the corpus); the **slot is a
global**. In our run every such global resolves to **0**, so the background (`BG030A`), event CGs, and
the effect spritesheet (`AE001H`, an 800×400 4×2 grid of blob frames) **all set-texture into slot 0**.
Objects live-reference their slot, so loading the effect **evicts** the BG → grey; and the effect is
drawn full-screen from slot 0 (its object `src=(0,0 800x600)`) → the whole sheet (blob grid) covers the
screen. ⇒ The layering failure is a **slot-assignment** problem. **CONFIRMED CAUSE (2026-07-08, live
`AGE_DIAG_SETTEX` trace):** every `set-texture` slot = `G[0x62452]`, written by `query-gfx-object?`
(`0x215`) which returns **-1** for the (correctly-unregistered) CG/effect handles → the fallback at
SC0000 `label_12649` does `G[0x62452] = lookup-array-2d(rec[s3]=G[0x3239], G[0x62450], 3, 0)` = **0**
because the slot table `rec[s3]`/`G[0x3239]` is **empty**. That table is filled by `call label_125bd`
(SC0000 `0x50f`, slots 4..11), which is reached **only through the scene-coroutine framework** — the
`G[0xaba5c]` gate (`0x450`) + op `0x140` (`u0041F9C0`, coroutine LABEL/yield `"LABEL" "J"` @ `0x46d`).
**Fixed by the bounded scene-coroutine host model:** `0x140` runs the setup body once, `label_125bd`
fills the eight slot records, and SC0000 resource `0x23` loads into assigned slot 5 instead of slot 0.
This was not a compositor/z-order/blend bug. Diagnostics: `AGE_DIAG_SETTEX=1` env → VM logs each `set-texture`
slot operand + `query-gfx-object?` result.
### Differential offset-path oracle — engine-vs-VM control-flow diff (2026-07-09)
**Method (lever #3 of the RE-front-loading program).** Run the same scene in the real engine and our C#
VM and diff the **executed script-offset path**. Both run the same bytecode, so the opcode at each offset
is static (from disasm); the first place the two offset sequences differ is exactly the branch/opcode/state
we modeled wrong — cheaper and higher-signal than diffing effects, and precisely where the render-drift
walk-backs lived. Tools: `tools/frida/trace_engine_ops.py` (engine capture) · `Age.Cli trace <SCENE>
--boot --trace-json` (VM capture) · `tools/diff_optrace.py` (align + first-divergence). Spec/plan:
`docs/superpowers/{specs,plans}/2026-07-09-differential-oracle*`.
**Capture method that WORKS = the operand hook `vm_operand_fetch@0x41b940`** (thiscall `ecx=ctx`; per op,
`offset=(pccodebase)/4` from `cur_ctx_index@0x53d14` / `frame_pc@0x53d2c` / `frame_codebase@0x53d28`,
per-context stride `0x78`). Validation: **100% of captured offsets land on valid SC0000 instruction
starts** — proves both the ctx-field offsets and the `(pccodebase)/4` math. **The tick hook
`adv_interpreter_tick@0x410fb0` does NOT work** — plain-JS `Interceptor.attach` there sees `ecx≠ctx`
(0 entries), so the recon gate's *tick* path is closed too, not just the CModule path noted under Frame
cadence. Two capture caveats, both handled:
- **Frida must hook BEFORE the scene loads.** The scene-entry setup (decl preamble + first CG load at
`0x802 mov G[0x62424]=0x23; call label_12649`) runs in a µs burst at load; a capture started mid-scene
floors at ~`0x80d` and misses it. `trace_engine_ops.py` writes `build/tracer-live.flag` once the hook is
installed → launch it in the background and gate the New-Game trigger on that flag.
- **Operand mode skips zero-operand ops** (stmt-begin/end markers, script-entry `0x259`) — they never
trigger an operand fetch. `diff_optrace.py` filters the VM trace to argc≥1 ops (same subsequence);
control flow is preserved (markers don't branch). `--full` disables it for a hypothetical tick capture.
**Codebase identification.** The engine trace tags each op with its script's `codebase` pointer (a per-run
heap address). `pick_scene_codebase` picks the codebase whose offset sequence shares the longest common
prefix with the VM trace. From a boot→opening capture (7 codebases / 265k ops), SC0000 = `0x09c1afe8`
(13252 ops, 100% valid, first offset `1` = offset 0 filtered).
**FIRST DIVERGENCE FOUND (the tool's first catch + self-test).** On the SC0000 opening the VM and engine
agree for **27 ops** (including the coroutine op `0x7b` @ `0x79`, which matched), then **diverge at offset
`0x8d` = op `0xa0` (jcc) on `global-int G[0x6c1]`:** the engine falls through to `0x94` (the op-`0x90`
hotspot-chrome registration block) ⇒ `G[0x6c1]≠0`; the VM jumps to `label_df` (`0xdf`) ⇒ `G[0x6c1]==0`.
Because both ran `1..0x8d` identically, `G[0x6c1]` is set **before** SC0000 — by pre-scene *system* boot
the VM's cold `--boot` (INITCONFIG/INIT2/INIT) doesn't replicate (the **two-boot gap**). `G[0x6c1]` is an
unlabeled but heavily-used scalar (766 uses) in the **same cluster as the op-`0x90` hotspot flags
`G[0x6c9..0x6cd]`** = ADV-chrome/input state. Headlessly benign (no input) but a genuine VM-side state gap
— and a clean demonstration that the oracle localizes a mis-modeled branch to a single instruction. It is
NOT the predicted coroutine yield (`0x140`/`~0x50f`); the oracle reports whatever diverges *first*, and it
surfaced an earlier state hole. **This is now the repeatable way to localize a mis-modeled op/state.**
Phase-2 extension (deferred): effects-diff (global-bank / gfx-registry writes) for branchy scenes.
### Scene-entry state snapshot — auto-seeding single-scene runs (2026-07-09)
**Problem the oracle surfaced:** single-scene VM runs diverge from the engine because they lack the
pre-scene global state the engine accumulates over `SYSTEM4 → … → TITLE → New Game`. `--boot` reproduces
only the data `*INIT` scripts; flags like `G[0x6c1]` (ADV-chrome enable) are set later and missed.
**Solution — `capture_global_writes.py`** hooks `vm_operand_write@0x425fb0` and logs every global-int
write as `(codebase, index, PLAINTEXT value)`. **Key: the helper receives the plaintext value** before the
engine encodes it into the obfuscated global store (rotate+XOR with the per-session cookie `ctx+0x55120`)
— which is exactly why the shelved flat-int32 scans (`global-memory-re.md`) found nothing, and why hooking
the WRITER is clean (no de-obfuscation). ABI: thiscall `ecx=ctx`, `[esp+4]`=operand index, `[esp+8]`=value;
the global index/type come from the instruction's operand slot (`framePc + idx*8`; `type=*(opnd-4)`,
`index=*opnd`), type 3 = global-int.
**Packer gotcha (solved):** AGE.EXE unpacks in-place at `0x400000`, so a `--spawn`-time hook hits packed
bytes → Frida "unable to intercept function at 00425FB0". Fix: poll `0x425fb0` until the real prologue
(`6aff 6836a85600 64a1…`) appears (unpack done), THEN attach. `--spawn` is required for completeness
(attach misses pre-attach boot writes); the tool also kills the spawned pid on setup failure so a JS error
can't leave a suspended windowless orphan. (First bug hit: `SIG` hex without `0x` → JS `create_script`
SyntaxError → resume never ran → orphaned suspended game.)
**Validated:** a real boot→New-Game→SC0000 capture (34,008 globals incl. `G[0x6c1]=1`, `G[0x62424]=0x23`
the resId) loaded via `Age.Cli trace SC0000 --state <snap> --trace-json` seeds the VM to match the engine's
**entire opening** (542 ops, no non-realignable fork) with ZERO manual seeding — confirming the
pre-scene-state theory and giving a general auto-seed for single-scene fidelity. **Residual:** a 2-op color
detour (`0x202/0x203` @ `0x122d0`, writing `G[0x62451]`) the full state does NOT fix = a real branch/op
difference to chase (not state). **Caveats:** snapshot is playthrough-specific (best for canonical entry
points — new-game opening, chapter starts); v1 captures global-INTS only (type 3; strings/floats TODO);
includes the scene's own early writes (can exclude by codebase for a pure pre-scene boundary).
---
## Native walls backlog (targets for this loop)
- ~~**call-script dispatch**~~ — **SOLVED** (above): `call-script <id>` = raw SYS4INI file index.
- **decision→scene** — how `0x62ccf`/the decision selects the next `SCxxxx`. Now narrower: scenes load
via `call-script`/the same SYS4INI-index loader, so the open question is only where the decision
value is turned into a scene *id* (a caller of SCJUMP; re-aimed away from `u00428010`).
- **op `0x60`** (`u0041A270`) — the rand-like value gating 1732/1755 SCJUMP decisions.
- ~~**gfx command-buffer**~~ — **DONE** (the `0x2120x21a` positioned-object subsystem = the rendering
drift): all 14 ops reversed + implemented against a host-side `GfxState`, and the missing INIT2 boot
state supplied via `--boot`. CGs render (screenshot-confirmed). See the op `0x215` finding + "The render
drift's SECOND half" above. Remaining: `AE*` alpha/blend (deferred) and cold-object anchors.