Learning to play Sil-Q
Research report · September 8, 2026. Edited for clarity September 9; study results and the materials used by the agents are unchanged.
Astra at medium reasoning completed ten Sil-Q sessions, each with eight practice games and two scored games. It won 0 of 20 scored games. Two sessions at maximum reasoning were interrupted before scored play, leaving the comparison between reasoning levels unanswered.
Sil-Q asks the player to descend into Angband, take a Silmaril from Morgoth's crown, and escape carrying it. Skills, equipment, light, noise, and consumables shape the route down and the chances of surviving the return. A coding agent must also build a reliable way to read the terminal and act on it.
Our evaluation follows the repeated-practice structure of EBR-Bench. Conversation, notes, and agent-written controllers persist across the ten games in a session. Each game starts with fresh private randomness; each independent session starts without history from the others. The final score is the better of the two scored games.
Results
All ten medium sessions completed without reaching a resource limit. Of the 100 scheduled games, 99 ended in native death and one practice game was voluntarily retired. None escaped with a Silmaril. This was below the predeclared saturation threshold: a scored win in at least nine of ten sessions would have required us to choose a different game. Gate record.
| Condition | Sessions included in results | Scored wins | Sessions with a scored win |
|---|---|---|---|
| Astra medium, 8 practice + 2 scored | 10/10 | 0/20 | 0/10 |
| Astra max, 8 practice + 2 scored | 0/2; both excluded | No scored games started | No result |
| Astra max, fresh control | Canceled after scope reduction | No result | No result |
The exact 95% interval for the probability of a session having at least one scored win is 0–30.85%. With only ten sessions, observing no wins leaves substantial uncertainty. An all-zero bootstrap interval would hide that uncertainty.

The mean maximum depth in scored games was 250 feet, with a session bootstrap 95% interval of 187.5–317.5 feet. The deepest scored game reached 500 feet. Depth helps describe play, but earns no partial victory credit: reaching the bottom is only part of the objective, and a fall can increase depth after a tactical error.

Depth varied across game positions and rose in the final two. Without the fresh controls, this does not estimate the effect of practice. The median game lasted 369 seconds and the longest 11,180 seconds, below the twelve-hour limit. The medium sessions used 4,390 model requests at an estimated cost of $2,094.75. Results and session reviews.
Why the max comparison is incomplete
After the ten medium sessions finished, the user reduced the follow-up to the two main max sessions already launched, selected by dispatch order. The two fresh controls were canceled. No further sessions or replacements were launched. Dated scope amendment.
Max session 01 was stopped at 23:00 UTC on September 8 following a request to wrap up, after approximately 14.5 hours. Seven practice games had ended in death; the eighth was alive at 400 feet with 49 HP. Neither scored game had started. Review covered all 1,018 tool calls, 217 visible assistant messages, and 88 available reasoning summaries. All 19,540 native frames replayed exactly, including the interrupted game. The final canceled request has unknown usage; there were no earlier API errors. Session 01 review.
Max session 02 recorded two API timeouts after its third practice game. The study's admission rule excludes sessions with model errors, so its retry loop was stopped at 13:12 UTC on September 8. All 264 completed tool calls and 7,343 native frames were reviewed; the replay was exact. Session 02 review.
Both attempts remain in the session count, with their exclusion reasons. Neither the timeouts nor the interrupted games count as game losses. Their request records verify Astra max, the 128,000-token output allowance, and the effective 997,500-token context window. They provide practice traces, but no scored max result or estimate of a reasoning-level effect.
Why Sil-Q
Our first candidate was Race for the Galaxy. After correcting a rule description and repeating the experiment, Astra medium won in nine of ten sessions—fourteen wins across twenty scored games. That met the saturation threshold, so we kept the result and chose another game. The earlier study records that decision.
Sil-Q combines character building, partially observed exploration, and a clear retrieval-and-return objective. It shares Brogue's genre, so it adds less variety than a strategy game would. Its explicit skills and abilities offer different choices, and the native terminal lets us retain the game's own displays and rule descriptions. A Wesnoth campaign was another candidate, but its scenario transitions, persistent units, gold carryover, and fog of war would have required a broader adapter. Selection record.
Sil-Q was absent from the ten-game SpeedrunBench list checked on September 7, 2026. This was a check of that dated list, not a claim that the game has never appeared in another benchmark.
Interface, scoring, and validation
The agent receives the native 80×27 terminal: characters, colors, cursor visibility, and warning bells. It can open inventory, equipment, character, ability, and help screens. The engine and scorer run in a separate protected service; the agent receives no hidden map or monster table.
Each key advances the engine until it next waits for input. Resting may advance many turns, and dismissing a message may resume combat already in progress. For a batch of keys, the client saves every intermediate screen and prints the last one. Agents can also write controllers against the public HTTP interface.
A win requires a native escape at the surface with at least one Silmaril in the
final inventory. The grader counts quest items and checks the native escape
flags. Native escape sets both escaped and the death flag; rejecting every
record with the death flag would reject real wins. Tests cover successful
escape, death with a jewel, empty-handed escape, crown variants, and empty
inventory stacks.
Native save/load, wizard entry, spoilers, configuration changes, and replacement games are disabled. A protected journal records inputs, observations, and endings. Replays compare those records with the run's frozen engine. These checks cover specific scoring and isolation failures; native C bugs and container vulnerabilities remain possible. Methods.
Defects found during development
GPT-5.5 medium pilots exposed clipped batch output, a validator that rejected a consumed potion before its inventory slot was cleared, and an inaccurate rest-command description. We fixed each and repeated validation. A later check found that API context metadata had not configured the smaller client context window. The affected research cohort, v1, was excluded.
Review of v2 found a display defect: at 24 rows, the second song and depth indicator occupied the same row. Starting or stopping a song could erase depth. One agent continued describing its position as 300 feet after descending to 350 feet while that field was blank. All 36,610 frames replayed exactly—the replay faithfully reproduced a faulty layout.
Native fixtures reproduced the collision. A 27-row terminal preserved both song lines and both depth fields. We excluded v2, changed the geometry, and repeated the development checks before launching v3.

The corrected build passed 49 engineering and admission checks. Sanitizer tests covered 5,823 native frames, including 4,758 also reproduced by the production engine. A fresh GPT-5.5 medium pilot died to a White Wolf at 250 feet. All 506 frames replayed exactly; 744 captured public views covered every frame. Review included 258 requests and 264 tool results, with no resource limit reached, clipping, or compaction. Pilot review.
The development agent also played through the public interface, testing combat, abilities, doors, traps, recovery, scrolling, and song changes. It deliberately quit at 100 feet with full health. All 270 frames replayed exactly and all 92 recorded public screens matched. The developer had prior source knowledge; this was an interface check, not an independent human baseline or a full-game victory. Direct-play record.
What agents changed during play
The admitted medium sessions contain concrete attempts to learn from mistakes, alongside recurring controller failures.
In bundle 03, a character found Shadows through the native monster list even though their map glyphs were blank. It later rested fourteen times while another Shadow attacked because the controller lost warnings in intermediate messages. The agent added a message buffer before the next game, but its danger check still omitted “tears at you.” It then rested under a Nightthorn until death. Shadow encounter · Next-game resting loop.
Bundle 07's controller classified monsters by alphabetic glyphs and missed
hostile human @ characters. Its final character rested fifteen times under
Easterling attacks, drinking a healing potion along the way. The native screens
and messages contained the warnings.
Final encounter.
There were successful recoveries too. Bundle 08's final character survived an Orc fight at 8 HP, killed two soldiers, withdrew, and recovered fully before descending to 250 feet. It died later among a new group of enemies. Recovery.
Bundle 10 survived repeated crises: recovery from 8 HP at 350 feet, two retreats from 400 feet at 19 HP and 7 HP, and a brief escape at 1 HP. Strength drain later reduced damage and carrying capacity. Excess equipment stopped movement, leading to 402 rejected movement inputs at one native turn; dropping items restored mobility. The game ended in death at 400 feet. First rescue · Late retreat.
Agents' explanations sometimes disagreed with the trace. Bundle 04 blamed an Orc charge for a large health loss after a weapon switch removed its shield; the recorded damage came from arrows. The charge happened later, after the shield was restored. Bundle 10 blamed a navigation error on Easterlings sharing the player's color, but the inspected screen showed distinct colors. Arrow damage · Later charge · Bundle 10 review.
The interrupted max game
Session 01's eighth practice game offers further examples, though it contributes no capability result.
At a forge, the agent made a shield, helm, and boots after fixing a helper that expected “Boots” instead of “Pair of Boots.” All 388 frames of that sequence showed 49 HP. Later plans for Boots of Speed and gloves granting freedom of movement remained previews away from a forge, with no uses available; those items were never made. Forge sequence.
When surrounded by soldiers and trolls at 300 feet, the agent bought Elbereth, sang, and escaped as enemies fled. It added fear management to its controller. Later it used the same song while trapped in Attercop webs, watched two spiders flee, and broke the second web after seven movement attempts. The 72 frames covering the web encounters and escape retained 49 HP. The next revision added Attercops to its fear targets and stopped on webs. Combat damage can also cause fear, so these observations do not isolate the song's benefit. First escape · Web response.
The agent also fought a mountain troll with a chasm behind it, despite having read about knockback. A westward attack triggered a fall from 300 to 400 feet, leaving 30 HP and bleeding. The next four inputs acknowledged messages rather than choosing new moves. The character recovered, but its greater depth came from a tactical failure. Knockback and fall.
A Rage herb helped kill two Easterling warriors, but its red enemy colors confused the player selector. The agent inspected the colors and fixed that selector before continuing. Rage and repair.
Limits of the result
The win metric is at its floor for medium, and max has no scored result. The observed repairs show that agents changed their behavior during play. Without completed controls, they do not establish a causal learning effect or a benefit from higher reasoning effort. Prior knowledge of this public game is uncontrolled. Controller construction and debugging are part of the task, so these results characterize this coding-agent setup rather than planning ability alone.
The original manual contained material errors: Elbereth costs one Voice per turn rather than one-third, Mastery costs two rather than one, and Strength in Adversity also raises Dexterity. Both max agents noticed Elbereth's higher consumption. We cannot rule out an effect of the advice on strategy or survival. The measured sessions kept their original materials; the current guide corrects them for future use. Documentation errata.
Native wording and displays also have limits. Territorial monsters stop pursuit when line of sight is lost, a condition omitted by recall. A target sidebar can retain an earlier alertness label even when the newly opened monster list is current. Exact replay preserves these quirks. Pinned pursuit rule.
Public-view coverage was incomplete in two sessions. Bundle 02 lost one complete client view when the agent killed its helper between input and logging. Bundle 09's early logger omitted 2,016 complete views after accounting for numbered displays. Their protected native records are complete and replay exactly, but we cannot claim the agent read a missing view.
Each game allowed twelve hours and 100,000 keys, with a conservative $1,200 session spend limit. These were set before research; shallow losing pilots do not show that a winning trajectory fits. Incomplete, invalid, or resource-bound sessions cannot establish that the task is unsaturated under this protocol. Resource allowances.
Matching initial random streams does not ensure identical dungeons: builds and actions consume different draws. There is no independent expert completion baseline through this interface, and elapsed agent time is not a human task horizon. Reviews and direct play were performed by the development agent.
Cost and artifacts
The known request-cost subtotal is $3,308.55:
| Work | Estimated cost |
|---|---|
| Completed medium research | $2,094.75 |
| Five completed development pilots | $104.23 |
| Interrupted development check, recorded usage | $0.81 |
| Two earlier excluded research launches, recorded usage | $193.09 |
| Two canceled fresh controls, recorded usage | $23.15 |
| Max session 02, stopped after API errors | $119.69 |
| Max session 01, stopped for wrap-up | $772.83 |
Each canceled control has one request without usage; the API-failed max session has three and the max session stopped for wrap-up has one. Their missing costs are unknown. These request-based estimates exclude developer inference and hosting and are not invoices. Run and cost ledger.
The evaluation site contains the report and native replays. Methods, the development ledger, and the September 9 harness audit document the measured setup and subsequent changes. The completed study supplies a medium baseline and examples of controller adaptation; the max comparison remains open.