Development history
Entries below describe the state of the work at each stage. Statements about pending reviews or launches are historical. The research report gives the final results. The README describes the current harness; this ledger ends with the September 9 audit. Counts, run identifiers, and evidence references are retained.
2026-09-07: selected after Race for the Galaxy's completed saturation gate. Initial source inspection found debug commands, wizard mode, editable options/preferences and game-managed save/load flows in upstream. The production interface needed to disable those capabilities. The native terminal is the planned public observation source; score and private randomness remain in the protected service. No pilot or research run had started.
Before first GPT-5.5 pilot
Implemented native terminal capture, protected single-use trial lifecycle, strict key/frame contract, hash-chained journals and independent inventory/escape cross-checking. Fixed an upstream birth-menu invalid-letter bounds read. Native save/load/configuration/debug capabilities are blocked; normal gameplay remains native. The native September calendar is fixed; gameplay randomness uses a secret-key HMAC stream with separately seeded quick/cosmetic randomness.
24 tests passed, covering native positive and negative quest fixtures, exact replay, input no-ops, menu/observation RNG invariance, no-reset semantics, batch limits, timeout, secret filtering, tampering and infrastructure failures. Sanitizers covered 5,823 frames including forced test-only deeper screens at 100, 250, 500, 750, 950 and 1000 ft. The non-injected production cases replayed 4,758 frames exactly. Across 40 paired initial gameplay screens, 581 unseen monsters had their health changed without a public-screen difference; the full paired prefixes also compared 1,480 frames including unchanged character creation. These tests cover the exercised cases; other native bugs may remain.
The client/service containers passed a transport and isolation probe, including a stale request, attempted local score forgery, protected-journal audit, no game files/state mounts in the client, no exposed ports, internal-only networking, dropped capabilities and a read-only service root. The only initial unit-test failure incorrectly counted upstream .gitignore as a saved game; the corrected assertion compares file contents before and after blocked commands.
The first development pilot is GPT-5.5 medium, 2+1, with provisional 7,200 seconds and 100,000 input keys per game and a $300 conservative session spend guard. These allowances are provisional: the official manual estimates roughly ten hours for a human winning game, while programmatic input may be much faster. Resource usage, controller behavior and any binding caps must be reviewed before freezing a research protocol. No Astra request is authorized by this engineering pass alone; pilot review and direct play were still required before medium research.
Final pre-dispatch corrections and pilot start
The first launch preflight stopped before creating a session or sending a model request because the imported task lacked an Inspect discovery wrapper. Added the wrapper and a fresh-process discovery/contract regression. Also preserved native warning bells in public frames (invalid keys can leave text unchanged) and disabled player-name-specific preference loading. The rebuilt adapter passed 26 tests, including both new regressions, and repeated the full sanitizer/replay/hidden-state suite with the same coverage counts. The actual container probe passed again.
dev01-gpt55-medium then dispatched successfully at 2+1 with the provisional
limits above. The initial 55 completed raw API events confirm GPT-5.5 and medium
reasoning; full-run identity review remains required. An additional no-model
Inspect/client/scorer exercise passed both fresh trials, all scorers and
isolation checks. A separate 65-frame engineering game replayed exactly using
its frozen service image in a network-disabled container. The developer pointed
the model-log reviewer at that no-model replay artifact once; it correctly
refused because no Inspect model log exists. It made no model request or
admission decision.
First pilot excluded; second pilot started
The completed log for dev01-gpt55-medium contains 148 verified GPT-5.5 medium
requests, 154 tool calls/results, 119 visible assistant messages and 67
available reasoning summaries. The private review is summarized in
analysis/dev01_review.json. Only the first practice game started and none
finished, so this pilot was excluded.
Review found two defects. A 25-key birth batch exceeded the coding tool's return
cap despite a requested 50,000-token allowance. The CLI now saves every public
batch frame to an agent-owned history file before returning one final frame and
a range receipt; silq history retrieves any saved frame without advancing the
game. Direct HTTP still returns every frame. The second defect was a validator
rejection during ordinary potion consumption: native code decrements the last
item to zero, prints a description that can wait at a More prompt, then removes
the slot. Validation incorrectly treated that transient slot as corrupt. It now
accepts the native transition while independently requiring positive item
quantity for quest possession credit. Invalid raw frames are now journaled
before rejection, to help diagnose failures.
All 474 recorded native frames replayed exactly, recovering the previously unlogged failure at frame 475: the Slowness potion had quantity zero while its effect message awaited input. The model's final answer blamed its valid command and reported zero; this is evaluator-caused invalid evidence, not a model loss. Eight extra Space keys were developer continuation only, excluded from performance.
The original untagged container image became unavailable during development; its native executable and entire data tree had already been extracted and retained. Both matched the replacement runtime byte for byte, allowing the exact native recovery above. The original container image itself was not recovered. Future launches now retain immutable image tags and verified compressed docker-save archives before dispatch. Both archives passed a load-and-identity round trip.
The corrected implementation passed 35 tests. The new sanitizer-backed potion fixture initially failed to reproduce the boundary because its injected inventory was unsorted; fixing the test-only ordering reproduced the actual transition. A subsequent local extraction briefly used an older test-image tag; the final test evidence uses the corrected build. Neither issue affected production runs. An isolated client/service probe recovered all 25 intermediate batch frames exactly from local history and matched every archived screen to the protected journal. The no-model Inspect/client/scorer check also passed again.
dev02-gpt55-medium started fresh with the same provisional 2+1 schedule and
resource allowances. Its initial four raw API records confirm GPT-5.5 medium.
Its complete review, direct public-interface play and resource calibration
remain necessary before Astra research. No Astra session had started.
## Complete second pilot and direct play (2026-09-07)
The second GPT-5.5 medium pilot completed with 189 verified requests, 195 complete tool calls/results, 183 visible assistant messages and 76 available reasoning summaries. All were reviewed; 769 native frames replayed exactly. The public observation audit matched 941 views across 758 distinct frames without a mismatch. Eleven late intermediate frames fell after the final auxiliary workspace snapshot; the CLI reported saving them, and the final frame was printed correctly. There were no API/service errors, rejected inputs, missing tool results, compactions or clipped outputs. Estimated request-priced cost was $14.633331, not an invoice.
All three games were voluntarily retired: 100 ft after 972.09 seconds, then 50 ft after 176.82 seconds and 50 ft after 94.21 seconds. These were allowed zeros under the original contract, with no resource cap reached. The agent learned native ability prerequisites, used Elbereth successfully, defeated early enemies and rested to full, but repeatedly misread room coordinates and stopped exploration early. These retirements record a choice to stop, not an unavoidable defeat.
During review we corrected the guide's rest-command entry: this pinned native
version does not ask for a duration; it rests until recovery or interruption.
The agent had recognized this behavior and used it successfully. Nevertheless,
this older-guide pilot is superseded for readiness, with its retirement outcomes
retained as development evidence. The fresh dev03-gpt55-medium pilot uses the
corrected guide, the same 2+1 development schedule, 7200 seconds/game, 100000
keys/game and conservative $300 session guard. Its first ten actual API records
confirmed GPT-5.5 medium.
The first direct public-interface play used a Noldor/Fingolfin melee character,
selected skills and abilities, traversed locked doors and scrolling corridors,
triggered a slow trap, fought Mewlips, used Elbereth fear, picked up a spear,
and descended twice to 150 ft. Native resting recovered health and was
interrupted. The developer then deliberately quit through Q/y/@; native grading
correctly recorded zero. All 181 frames replayed exactly. This was
development-agent validation, with prior source knowledge, not an independent
human baseline or a claim to have won the game. No hidden state from that run
was read while choosing actions. analysis/direct_play01.json and the native
replay preserve the evidence.
A separate public continuation investigated pilot 2's claimed dead pocket. It recreated all 149 public frames of that retired trial exactly before any new decision, with the corrected guide. Twelve additional public keys had reached the room north-west through the visible corridor and door. The original pilot outcome remains unchanged; the continuation was a developer check.
The research pipeline now refuses Astra dispatch without a registered cohort, full GPT-5.5 review, direct-play replay and resource calibration. It freezes ten independent matched stream schedules for medium, max and fresh max control in advance. Max additionally requires all ten medium sessions to be reviewed, non-resource-bound, and below the prospectively stated 9/10 saturation threshold. A stale evidence hash, changed game, altered stream schedule or forged saturation count is rejected. An initial analysis-only check incorrectly required every ending to be native; it was corrected before any Astra request to retain allowed voluntary retirement as zero while distinguishing it from native death. Seven research-gate checks pass. The latest full suite passed 41 checks before the new retirement regression; the resulting suite has 42 checks.
The public replay reader passed a browser check of every text cell, color attribute, palette value and cursor in all 475 frames of the recovered first pilot, with no JavaScript errors, broken links or mobile page overflow. More replays are being added as their reviews finish. No Astra session had started.
## Final direct-play checks and admission review safeguards
The corrected-guide continuation ended after 149 original public frames were
reproduced exactly, followed by 34 developer keys, armour collection/equipment,
a second room and automatic rest, then deliberate native quit. All 183 native
frames replayed exactly. The original retirement was unchanged. The main
analysis/direct_play.json binds the corrected implementation to both direct
runs and their full replay/public-observation evidence.
The admission check now requires every security-pattern finding to be reviewed with evidence. Blocked attempts and benign regex matches can be admitted, so the filter cannot selectively remove legitimate losses. A successful bypass or an unresolved finding prevents admission. Eight research-gate tests and all 43 engineering tests passed (10.45 seconds for the full suite).
The browser checked all 1608 frames of the six currently exported replays, including every text cell, attribute, native palette value and cursor: no JavaScript errors, broken links or mobile page overflow. The bounded public export audit checked 76 text files and all 1608 frames against 17 retained private stream keys and common credential/encrypted-payload patterns, with no findings. The first version of that analysis-only field allowlist mistakenly expected a public native-event kind and omitted cursor visibility; it was corrected to the actual strict projection schema before the successful audit.
Corrected pilot admitted and resources frozen
dev03-gpt55-medium completed successfully: learning retirements at 50 and 100
feet, followed by native scored death to a Mewlip at 100 feet. Every one of
11082 native frames replayed exactly. All 435 raw requests and responses
verified GPT-5.5 medium, with all 454 tool results, 363 visible assistant
messages and 190 available reasoning summaries reviewed. Two context compactions
occurred. There were no service errors, boundary findings, missing usage records
or resource-bound games. Estimated request-priced cost was $34.377892.
Three workspace controllers and their revisions were inspected in full. All 4952 saved controller records matched native public values and executed keys; all 3425 complete printed controller records matched those logs. Three tool outputs clipped repetitive logging, with the full retained logs and native journals supplying the omitted evidence. The new review script initially split quoted space-containing action strings incorrectly; correcting that analysis parser yielded an exact comparison. No engine or client change resulted.
The controllers repeatedly confused compressed map cells with movement cells, failed to remember explored frontiers, or waited when dark. The last controller overgeneralized an up-stair maze event, then oscillated at 100 feet while scouts shouted. A Mewlip extinguished its fading torch; it waited through unseen attacks until native death. These are agent strategy/controller failures. The repaired single-key response KeyError and shell quoting error were also agent code errors. Retirements remain voluntary zeroes, distinct from unavoidable native defeats.
Full screen payloads were independently captured for 940 unique frames, with 1354 matching views and no differences. Most controller-only screens were not printed in full; their logged projections, the tested public serialization and all-frame native replay provide the additional evidence. The review does not claim independent capture of every HTTP response body.
RESOURCE_CALIBRATION.md freezes 12 hours and 100000 keys per game and a $1200
conservative session guard before Astra. All ten independent medium schedules,
their matched max schedules and fresh max controls will be generated before
outcomes. Full first-session review precedes the remaining medium dispatch.
Client context mismatch caught before a scored Astra result
The first research session made 41 Astra medium requests (40 completed, one cancelled) before live client telemetry exposed a 258400-token effective window. Inspect model metadata reported 1050000, which had not configured the Codex client. All of research v1, including its unlaunched schedules, was excluded. Its sole partial game had 680 exact native frames, no completed trial and no compaction. It is not a saturation result or a context-caused loss.
The next GPT-5.5 check explicitly set context and output allowances, but still reported 258400: the bundled model catalog imposed a 272000-token maximum. That check, dev04, was cancelled after 23 requests (22 completed) and 67 native frames, with no completed game. Both partial journals replay exactly and remain retained with their exclusion reasons. Missing cancelled-request usage is not filled with zero.
The final correction builds a custom catalog from the pinned client's own bundled catalog, changing only three context fields for GPT-5.5 and Astra. The image preflight verifies every other field is unchanged and runs the actual client against the configured catalog. Live dev05 telemetry now shows the expected 997500-token effective window and every checked request allows 128000 output tokens. Admission requires both observations. Forty-six engineering checks pass. See resource calibration for the client configuration and upstream references. Direct public play is repeated on this build, followed by full dev05 audit and a fresh cohort freeze.
Corrected scaffold admitted for fresh Astra research
Dev05 completed its single scored game with native death to an Orc soldier at 200 feet while poisoned. It took 959.93 seconds, 412 keys and an estimated $9.986348. All 159 GPT-5.5 medium requests verified both identity and 128000-token output allowance; live telemetry showed 997500 effective context. Peak input was 179950 with no compaction. All 171 tools, 103 visible messages and 64 available summaries were read; 413 of 413 native frames replayed exactly and were independently captured as public screens, with 560 matching views. Eleven public-history coordinate/ruler outputs matched exactly. A repaired missing import was the only agent execution error; no service error, clipped output, missing usage or boundary finding remained.
The agent navigated manually to 200 feet, killed two Orc warriors, rested to full health and equipped gauntlets. It then entered a brood-spider room in a long movement batch and became surrounded. Emergency skill purchases worked without advancing turns, but poison and pending monster attacks ended the game. Native More prompts can precede a status redraw: displayed HP lagged protected HP during combat resolution, exactly as the native replay reproduces. Clearing such a prompt can resume already-pending attacks. These are native terminal semantics, not service-created actions.
Direct play on the same build completed another native trajectory, all 81 frames exact, with doors, item pickup, descent, rest and native quit. Current readiness binds this evidence and 46 passing checks. Fresh research v2 uses new private streams; no v1 stream or partial outcome is reused.
Auxiliary collector coverage during research v2
The first v2 Astra session is running. A read-only collector review found that files larger than 40 MB were skipped before the omission-report branch; a large Codex session log could also lose its context/compaction metadata from the snapshot. The collector now lists oversized eligible workspace files as omitted and streams session logs line by line for resource metadata. A regression test uses both a >40 MB workspace file and >40 MB session log, verifying the omission record, final context/compaction fields and exclusion of encrypted content. It passed independently of the 46 checks frozen at readiness.
The watcher was restarted with this change. No game/client/provider code, guide, resource contract, random stream or agent observation changed. The original frozen run sources remain intact. Public/controller response coverage is still reported per run: an omitted file is not silently counted as captured.
The auxiliary capture ceiling was then raised from 10 MB per file / 40 MB total
to 100 MB per file / 200 MB total after the first Astra controller history
exceeded the old ceiling. Oversized files are still explicitly reported. This
retains more agent-authored evidence without changing agent inputs or the frozen
game/scaffold. The updated oversized-file and >40 MB session-telemetry test
passed (1 test, 0.27 s; .private/pytest-workspace-capture-v2.log).
The research article is now rendered from BLOG_POST.md by
scripts/build_site.py, with the pending status explicit. An article-only
browser check passed desktop/mobile rendering and every local link
(analysis/article_qa.json); the replay renderer was unchanged.
scripts/plot_results.py produces standalone PNG/SVG/PDF figures only for
complete admitted conditions and currently emits no performance figure. Binary
probability plots use exact binomial intervals to retain finite-sample
uncertainty at zero wins. The public run ledger now includes excluded launches
and separate known subtotals for interrupted requests; it does not treat missing
usage or an active run as free. These changes affect reporting/auditing only.
Frozen readiness was reverified unchanged.
The first research controller later accumulated 90 MB of full public screen history. The read-only auxiliary snapshot ceilings were raised to 1 GB per eligible file and 2 GB total, with every omission still explicit. Prior checkpoints remain retained. The sparse oversize-file/large-session-telemetry test was updated to exercise that actual boundary; this changes no game, model, or resource contract.
During the first research session, auxiliary audit checks were extended to validate full controller screen dictionaries and trial phase, status, key counts, and bounded remaining-time fields. Mixed outer tool calls can contain both menu output (frame, message, footer) and movement/retreat output (frame, message); the checker now recognizes those actual projections. Every accepted projection still matches native fields exactly. Plain printed coordinate/list text retains its brackets in the compact review renderer. A separate live-to-final continuity checker matched all 171 tools, 103 visible messages and 64 available summaries on the completed final pilot. These are additional review tools, not changes to game/scaffold or the frozen readiness evidence.
Replay export now checks the full evidence-bound admission record before labeling a session admitted; excluded evidence requires completed review, and direct play requires exact full native replay. Replay pages support native frame links. Browser checks cover direct navigation, stepping, sharing, reload, invalid-frame fallback and mobile overflow; the initial invalid same-page link check exposed and fixed a hash-change fallback in the presentation layer. All game/source and frozen readiness files remain unchanged.
Additional first-session review checks verified the printed Orc captain color and enemy positions in scored game 9. An initial checker included the top message line when collecting map glyphs; restricting it to native map rows resolved that false alarm. All 35,790 captured controller screens and their phase/status/key metadata matched the protected journal at this checkpoint. This changed audit code only, with no game/scaffold changes.
Native geometry correction after v2 rollout review: the pinned upstream
prt_song clears zero-based rows 21/22, while ROW_DEPTH is terminal height
minus 2. At 80x24 the fields overlap. In v2 scored game 9 the controller
descended to 350 ft at frame 4136, later saw blank depth, and continued
describing 300 ft until another redraw. Separate test-only native fixtures
reproduced blank depth after an empty-song redraw and replacement by a second
song; 27 rows preserved both songs, current depth and minimum depth. The whole
v2 cohort was excluded and operator-cancelled before a scored ending. All 100
streams are retired.
The excluded first session has 490 requests (489 completed, one operator-cancelled with unknown usage), 489 fully reviewed tool results, 33 visible assistant texts and 59 available summaries. All 36,610 native frames replay exactly; all 36,610 unique public screens were independently captured. The auxiliary controller check matched 36,599 raw screens/metadata and 589 numbered screens. Known request cost is $190.344817; it is not a complete invoice or a capability result. All source/images, audit evidence, eight native practice deaths and the unfinished scored game remain retained. The final native frame at 350 ft has 41 HP; it is not recorded as a scored loss. The test demonstrates why exact native replay alone cannot establish that a native interface is sound.
Production geometry is now 80x27, with strict matching public-schema dimensions.
Two native song/depth regression cases were added. The new build passed all 49
checks in 10.86 seconds, 5,823 sanitized frames, 4,758 exact production replays
and 40 hidden-state pairs (581 hidden monsters changed, no public difference).
Historical replays keep their original geometry; the viewer and public-output
audit now use each frame's declared dimensions, while compact tool review reads
the frozen run's native height. Fresh topology/Inspect checks, GPT-5.5 medium
validation and direct play precede fresh v3 research. V2
methods/readiness/source were copied to .private/scaffold-v2-freeze before
replacement.
## Corrected 27-row validation complete
Dev06 completed one native White Wolf death at 250 feet, HP -4, 411 player turns, 505 keys and 1771.01 seconds. All 506 native frames replay exactly; 744 complete public views cover every frame, and 15 custom coordinate outputs reconstruct exactly without executing the agent's code. All 258 GPT-5.5 medium requests, 264 tools/results, 244 visible texts and 97 available summaries were read. All requests used the expected model identity and 128000 output allowance, with 997500 effective client context, peak input 318161 and no compaction. No missing usage, service/model error, clipping, execution failure or resource cap remained. Known request-priced cost is $35.166702. Final pending-combat message prompts were checked against native turn/HP evidence; they reproduce native mid-action displays, distinct from the earlier persistent song/depth collision. The complete review and evidence hashes admitted dev06 as development evidence.
The development agent's direct-play04 used only the public interface until native quit and cleanup. It reached 100 feet, killed six Wolves and a Grimhawk, recovered from 21/41 HP, bashed locked doors, recovered from stun, bypassed plants and sleeping wolves, picked up a torch and inspected inventory, equipment and the compressed map. Elbereth start/stop at 50 and 100 feet retained both depth fields. Wall collisions and a declined trap confirmation remain visible in the full replay. All 270 native frames replay exactly and all 92 printed public screens match. The native quit ends at 100 feet with 41 HP and zero score (269 keys, 318 turns, 1063.07 s including developer deliberation). This is not an independent human win-rate check. The historical direct-play report is retained as analysis/direct_play_v2.json.
The live 27-row isolation probe and actual Inspect 1+1 birth/quit integration also passed. No game/scaffold changes were made during the fresh pilot or developer play. Readiness binds current identities, the full pilot and direct review, 49 passing checks, native fixtures, live topology, Inspect integration, methods and the unchanged resource contract before fresh v3 streams are generated.
Fresh v3 schedules were prepared after readiness verification: ten independent private streams per medium/max session, with the final two matched to fresh max controls. All 100 v3 keys are new. Only medium session 01 is dispatched initially; its full review precedes the remaining nine. The 9/10 saturation rule and conditional max gate are unchanged. A prepared or active schedule is not an admitted capability result.
Fresh v3 medium session 1 admitted
The first corrected-geometry session completed 8 practice and 2 scored games: all native deaths, no Silmarils, scored depths 150/150 feet and wins 0/0. All 168 Astra-medium requests, 167 tool results, 23 visible assistant messages and 51 available summaries were reviewed. Every one of 15,808 native frames replayed exactly; 15,967 captured full views covered all unique frames. The additional format audit checked 148 numbered screens/signals, 104 sampled native tuples and four coloured @ positions. The clipped initial manual read was reconstructed against frozen source; subsequent manual reads were exact. No model/service errors, missing usage, compaction or resource-bound endings occurred. Cost estimate: $19.989921, excluding developer work, hosting and invoice adjustments. The first session establishes no cohort saturation conclusion.
The controller's navigation, light, enemy classification and tactical failures
remain results. In particular, the seventh game's 300-foot maximum followed a
trap fall, and fixing player identification did not fix the alphabetic-only
enemy detector. The final game used Elbereth; all 805 selected normal-HUD frames
retained the correct depth. No game or scaffold changes were made during this
session. The completed qualitative record binds its evidence hashes and passes
admitted_run; the remaining nine medium bundles are dispatched with four
workers. They retain the pre-generated schedules and the same frozen game.
Fresh v3 medium bundle 04 admitted
Bundle 04 completed ten native deaths, with scored maximum depths 150/300 feet and no Silmarils or escapes. All 407 requests verify Astra medium; all 406 tool results, 42 visible messages and 46 available summaries were read. Independent replay matched all 25,763 frames, and captured public observations covered every frame. The two service rejections followed attempted actions after native death. Six agent code or shell failures were reviewed; there were no model errors, missing usage, compactions or resource-bound endings. Cost estimate: $117.309616. The session passed admission against the frozen cohort identity.
Additional review distinguishes actual visible terrain from the controller's remembered cells, and checks emitted actions, adjacent glyphs and colours against the journal. Corrupted glyph, colour, action and visible-cell counterexamples are rejected. High-health native screens abbreviate Health as Hth; the review helper now checks that alias, with six native-screen counterexamples passing. These changes affect review only. A developer stdout redirection mistake overlapped three replay JSON artifacts; their original bytes were archived and their exact equivalence to the complete reports proved before restoring valid JSON. See artifact repair record.
The final native action trace also corrects the rollout's own report: arrow hits caused the sharp health loss before the shield was restored; the recorded Orc charge came afterward. The eighth trial issued repeated waits beside an unseen Nightthorn. These are controller and reporting failures with intact native observations. Bundles 01 and 04 are admitted; no aggregate saturation conclusion or max run is authorized until all ten medium sessions pass admission.
V3 medium bundle 05 admission
Bundle 05 is admitted after all 385 actual Astra-medium requests, 384 tool results, 25 visible assistant texts and 47 available summaries were read. All 27,667 native frames replay exactly; 27,681 captured public views cover every frame. The session completed eight practice and two scored games, all native deaths, without a resource guard, retirement or replacement. The deepest practice character reached 400 feet; scored depths were 100 and 150 feet. Request-priced cost is $96.056586. The ten-session saturation gate remains pending, with bundles 01, 04 and 05 admitted.
Supplemental review checked six inventory dictionaries, visible colours and coordinates, 15 monster-awareness records, 76 song/voice history lines and 24 printed action traces, plus the existing 182 native-history projections. An original copy passed and four deliberately corrupted inventory, colour, awareness and position copies were rejected. The reviewer initially selected song casts and early-continue waits as printed debug steps; the range was corrected to the actual print sites. No game or model input changed.
Native records corrected working-note mistakes: trial 8's staff was of Mithrim; trial 9 wore studded leather; duplicate armour and cloak in trial 10's inventory did not mean its equipped slots were empty. The final potion showed no native healing message. The agent's assertion that an unidentified Emerald potion would heal was not accepted as a game fact. Every final Python program was statically read, backup variants compared by complete diffs, and final JSON/CSV results checked against the protected summary. No controller code ran on the host and no hints were sent into a research session.
The reviewer also added support for bundle 06's keys-before-status header and
colour highlights, and bundle 07's frames.jsonl archive. Initial live checks
match all 220 and 147 numbered views respectively. These are reviewer-only
changes; the frozen game, client, schedule and source identities are unchanged.
Research v3 medium bundle 07 admitted
All 348 Astra medium requests used the specified 128,000-token output allowance and 997,500-token effective context window. Review covered all 347 tool calls, 40 visible assistant messages and 48 available summaries. Peak input was 353,446 tokens, with no compaction, missing usage or model error. Five execution failures came from the agent's indentation, XP parsing, archive nesting and missing-coordinate handling; each was investigated and subsequently handled.
The independent native replay reproduced 22,624 frames exactly. Public captures covered every frame, with 22,639 matching views. Reviewer checks also covered 413 numbered screens, 20 signals, 303 history/coordinate projection lines and 300 ordered native message prefixes. The prefix checker rejects invented items, wrong HP, reordered pickups and incorrect frame bounds. Twelve additional counterexamples cover the scored projection formats. Twenty-nine clipped tool calls retain their original omissions; only source-matched ends were read. All sixteen final Python files and the report were read statically. The saved results JSON equals the protected summary. No agent code ran on the host.
Every game ended in native death. Scored depths were 200 and 350 feet, with no
jewels or escapes. The longest game took 2,375.57 seconds and 6,233 keys, below
all frozen guards. Estimated model cost was $92.571526. The first scored death
followed repeated resting while blind. The final controller ignored hostile
human @ symbols and rested fifteen times during its fatal Easterling fight.
Native observation checks preserved those enemies and attacks. The report's
claim of recovering from 3 HP in trial 6 was verified at frames 1066–1068; the
traces do not isolate the effect of the build. No developer hints, game changes
or outcome replacements occurred.
Research v3, medium bundle 03 admission
All 482 requests used Astra at medium reasoning with the frozen 997,500-token effective context and 128,000-token output allowance. Review covered all 481 tool calls, 57 visible messages and 49 available summaries. There were no model errors, missing usage records, compactions or resource cutoffs. Four agent code errors and one deliberate interruption of its equipment loop were retained.
All 43,759 native frames replayed exactly. The review matched 552 numbered screens and their attributes, 40,926 captured controller responses, 91 earlier history projection lines and 882 further native projection checks. Eight adversarial controls passed. Thirteen final Python files and two Markdown files were read statically; the results JSON equals the protected summary.
Three apparent HUD errors were native equipment-selection menus covering the health digits. Full native and adjacent restoring screens were inspected; the reviewer flags remain recorded with their resolutions. A lexical boundary flag matched the agent's own instruction to avoid hidden state and debug access. The full command used only the public interface. This paused queued dispatch; the finding was resolved explicitly before bundles 09 and 10 were resumed. No game or frozen implementation changed and no run was replaced.
All ten games ended in death; scored depths were 500 and 300 feet. The longest game lasted 10,315.07 seconds and used 27,085 keys. Request-priced cost was $282.301945. The first scored controller lost Shadow warnings when it cleared message prompts. The next buffered messages but omitted Nightthorn tearing attacks from its danger checks, then repeatedly rested. Several earlier successful Shadow fights and equipment improvements remain descriptive observations, without establishing a practice effect.
Research v3 medium bundle 08 admitted
All 204 requests used Astra medium with the frozen 128,000 output allowance and 997,500 effective client context. Peak input was 171,419 tokens; no compaction, model error, missing usage or boundary finding occurred. Review covered 203 tool calls, 24 visible assistant messages and 54 available summaries. The controller's single execution failure was a legitimate rejection after death; it corrected its message-clear loop to check trial status.
All 14,127 native frames replayed exactly. Captured public views cover every frame, and all 12,684 normal HUD checks match. The review verified 218 numbered screens, 49 supplemental native projections and nine adversarial controls. Two broad-checker flags arose from clipped output containing two screens in one call. Bounding each clipped block by the next header verified its retained ends; the original flags and the omitted text remain preserved. A buffered torch message was also checked against the active command's full interval, because an independent state read preceded collection of its earlier stdout.
The final workspace's 21 Python and two Markdown files were read, including complete diffs for birth variants. Its results JSON matches the protected summary. All ten games ended in death, with scored depths of 100 and 250 feet. The longest game took 543.36 seconds and the largest used 3,759 keys; total request-priced cost was $23.62463. No game reached a resource guard. Repeated song/rest loops, light exhaustion, movement under arrow fire and incomplete hazard handling remained controller failures. The final character's recovery from 8 HP is retained alongside its later fatal encounter.
Research v3 medium bundle 06 admitted
All 587 requests used Astra medium with 128,000 output tokens and the verified 997,500 effective client context. Peak input was 542,869 tokens. Review covered 586 tool calls, 75 visible assistant messages and 55 available summaries, with no compaction, model error, missing usage or boundary finding. All 28 final Python files and the play notes were read, including complete source diffs where a bulk display clipped. The result JSON and supplied source copies match.
All 29,431 native frames replayed exactly, and captured public views cover every frame. The 26,896 normal HUD checks have no discrepancies. The review checked 593 numbered screens, 486 highlight lists and 351 supplemental public projections. A mixed screen/entity-list clip in call 577 triggered the broad checker; both retained ends match the native projection. The original flag is preserved. Ten adversarial controls accept the original evidence and reject corrupted health, coordinates, frame labels, history, colors and depth.
All ten games ended in native death. The scored characters forged armour but died at 300 and 450 feet. The final character's transitions through 300 and 400 feet included a floor collapse and a troll-induced fall, correcting the model's account of planned navigation. It later recovered before the final encirclement. Unknown emergency consumables caused Slowness and fear; the last Warg attack killed it at frame 2622, with native terminal frame 2623. The longest game took 6,974.99 seconds; the largest key count was 9,760. No resource limit was reached. Estimated request cost was $280.492676. The effect of builds and tactics could not be isolated.
Research v3 medium bundle 02 admitted
All 741 requests used Astra medium with the frozen 128,000 output allowance and 997,500 effective context. Peak input was 898,747 tokens; one compaction occurred during the final game. Review covered all 739 tool calls, 95 visible assistant texts including the complete compaction summary, and 65 available summaries. All 30 final Python files and notes were read. No model error or missing usage occurred; seven agent code/command failures were reviewed.
All 36,432 native frames replayed exactly, with 32,215 matching normal HUD checks. The public audit covered 36,428 complete unique screens; three further screens match numbered displays. Trial 2, frame 7672 was absent from the agent's archive: call 155 killed its own auto.py helper between the native action and local append. The protected key and frame are intact, HP remains 41, and the next helper resumes the current game. This is retained as one unarchived controller view. It is not described as model-read or as an environment failure. The review additionally matched 734 numbered screens, 734 color lists and 270 supplemental projections. Seventeen controls accepted original evidence and rejected sixteen corruptions, including asynchronous progress depth and adjacent-door color.
Practice trial 6 was voluntarily retired after a routing stall; it consumes its scheduled zero. The other nine games ended in native death. The first scored character forged resistance rings, recovered from 3 HP and later died to an Attercop at 300 feet. The second died from poison at 100 feet before its planned forge. Neither game acquired a jewel or escaped. The longest game took 11,180.47 seconds and the largest key count was 10,695, below the guards. Estimated request cost was $599.399715. No production change or hint was sent to the running agents.
Research v3 medium bundle 09 admitted
All 541 requests used Astra medium, output allowance 128000 and effective context 997500; peak input 606938, with no compaction or model error. All 540 tool calls, 30 visible assistant messages, 62 available summaries and 35 final Python files were read. The two command failures were correctly rejected inputs after native deaths. Supplied sources and final result JSON match.
All 28,523 native frames reproduce exactly; 24,897 normal HUD checks match. The custom controller checker verified 668 numbered screens, 668 color lists and 26,416 archived HTTP views. Another 111 public projections match. Eight controller controls and 16 projection controls accepted original evidence and rejected 22 corruptions. The three clipped outputs preserve retained ends only.
The agent initially logged only batch responses. After accounting for full JSON and numbered views, 2,016 native frames in the first two games lack a captured complete client view. Call 77 fixed the agent's logger. Every protected native frame is retained and replayed; the audit does not invent missing client observations. No model or environment retry was launched.
All ten games ended in native death; the scored games reached 50 and 250 feet. The final archer recovered ammunition repeatedly, then exhausted it during an Orc fight. Golden Slowness and Mottled Terror preceded a final retreat; slowing returned at frame 1550 and an Orc warrior charged lethally at frame 1552. The longest game took 1647.59 seconds, largest key count 9920, and no resource limit was reached. Estimated request cost was $315.361155. Nine medium sessions are now admitted; the full gate remains pending until bundle 10 passes review.
Research v3 medium completed; max study started
Bundle 10 is admitted after review of all 527 Astra-medium requests, 526 tool results, 42 visible assistant texts and 55 available summaries. All 25,819 native frames replay exactly; 25,835 captured public views cover every frame. All 575 numbered screens, 3,671 supplemental projections and 22,111 normal HUD checks match. Thirty-nine controls accept original evidence and reject 37 alterations. Three clipped outputs retain explicit scope. All nine final Python files, their revisions and final notes were read. No model error, missing usage, compaction, retirement, boundary finding or resource guard occurred. Estimated request cost was $267.639850. Complete review.
Both scored games ended in native death, at 450 and 400 feet. The last game contained several recoveries and severe Strength drain. A burden loop repeated 402 inputs without advancing a native turn. Dropping equipment restored movement; a late 1-HP retreat survived briefly before a Grave Wight delivered the fatal blow. The review retains a controller return-branch bug, an unsupported color explanation, and the native territorial recall's omitted line-of-sight condition. A developer working-note mistake about Orcish Liquor was corrected before admission: it heals and removes fear but adds stun. These interpretation records are bound to the private review; no agent hint or game change was made.
The completed medium cohort has 0/20 scored wins and 0/10 sessions with a scored
win. The exact 95% session interval is 0–0.308497. There were 99 native deaths
and one voluntary practice retirement across the 100 games. The mean scored
maximum depth was 250 feet (session bootstrap 95% interval 187.5–317.5); depth
remains strictly diagnostic. All 4,390 medium requests cost an estimated
$2,094.747620. The reviewed medium gate was recorded as valid_unsaturated.
Astra-max bundles 01–02 and fresh max-control bundles 01–02 are dispatched with two workers each, preserving four total Sil-Q workers. Initial completed API requests in all four runs verify Astra max and the intended output/context allowances. Remaining bundles are not yet dispatched. Full max and control reviews, comparisons and the final article remain outstanding. Medium-only PNG/SVG/PDF figures were generated and visually inspected; pending conditions are not plotted as outcomes.
September 8, 2026 · User-requested sample reduction
The user requested one session, or at most three per reasoning level. All ten medium sessions were already complete and remain admitted. The max study now consists only of the two main sessions already dispatched (01 and 02). The two fresh controls were interrupted cleanly with SIGINT at 08:52 UTC, both during their first scored game; neither has a completed game or admitted outcome. Their 260 and 191 recorded native frames replay exactly. Recorded request costs are $10.272102 and $12.874852, with one interrupted request missing usage per control. No further sessions or replacements will launch.
The original cohort and medium gate remain frozen. A separate scope amendment records the user's instruction, the retained dispatch-order selection and the reduced claim. Launcher and dispatcher guards prevent further sessions. The report preserves canceled and unlaunched schedules as provenance, expects two max sessions, and permits a descriptive comparison only with their two matched medium bundles. Thirteen scope and research-gate tests passed, including rejection of extra launches, preservation of all medium results, exclusion of canceled controls from outcomes, and refusal to declare completion with an unreviewed max session. The game and running session contracts were not changed.
September 8, 2026 · Elbereth manual discrepancy
During the complete reading of max bundle 02's first practice game, its notes
identified a discrepancy between the bundled manual's one-third Voice per turn
and observed Elbereth consumption near one. Source review confirmed an
unconditional cost increment of one for Elbereth in the pinned sing()
function. The public count reflects actual Voice; the manual's rate is wrong for
this version. Max bundle 01 also exhausted Voice during its first game and
recorded the higher observed rate in its post-game notes. A causal effect of the
advice on either loss is unresolved.
Documentation errata and the article now disclose the error. The running game, guide and manual remain frozen. Native outcomes are retained as evidence under those actual materials, with the incorrect advice noted as a limitation. No hint was sent to an agent. A future corrected interface requires fresh validation and a distinct study identity.
Subsequent inspection confirmed that launch.py runs each task from its own
source snapshot, and task.py copies that snapshot's guide into the agent
workspace at initialization. The repository's current guide now corrects the
Elbereth rate and territorial-pursuit condition; the original guide is published
byte-for-byte as analysis/frozen_v3_game_guide.md. Neither running session was
modified. The revised documentation was not used in a model study or admitted
through a new readiness review. The version record binds both guides and the
verification evidence. No further model session was launched.
Study close and September 9 harness changes
Both main max sessions were interrupted before scored play and excluded: session 02 after API timeouts, and session 01 at the user's wrap-up request. The ten medium sessions remain admitted, with 0/20 scored wins. The scope amendment and report record the final reviews and costs.
Revision 0.4.0 separates victory accuracy from schedule completion, freezes final scoring under the action lock, and tightens deadline, journal, native-state, and seed validation. It also makes state-file replacement durable. The full suite passed 214 tests; two optional plotting tests were skipped. The audit records the fixes, smoke tests, historical journal checks, and remaining gaps.