Why we selected Sil-Q

We selected Sil-Q on September 7, 2026, after the corrected Race for the Galaxy study reached its predeclared saturation threshold: a scored win in nine of ten independent 8+2 sessions. We kept that result and chose another game rather than changing its difficulty or threshold. Prior gate.

Fit for repeated practice

Sil-Q requires the player to retrieve at least one Silmaril from Morgoth's crown and escape the Iron Hells carrying it. The native engine records escape and possession separately, allowing tests for death with a jewel, empty-handed escape, and success. Reaching the throne room alone does not win.

Attributes, skills, abilities, and equipment influence the entire descent and return. Combat, stealth, noise, light, and consumables create choices whose consequences may appear much later. Fresh private randomness changes each game, while conversation, notes, and controllers can carry experience between games. That makes Sil-Q a candidate for repeated-practice evaluation with a clear final objective.

At selection time, its difficulty for the models was unknown. The development checks and medium experiment were intended to test its suitability. Zero wins would also limit the measurement: depth and failures can describe play, but cannot be converted into partial victory credit.

Tradeoffs and alternatives

Sil-Q shares Brogue's genre, which limits the variety it adds. Its explicit character-building, ability, and stealth choices provide different decisions. A native terminal adapter also preserves the game's own descriptions, reducing the need to rewrite a large rules catalog—a source of error in the earlier task. We prioritized a complete objective and a testable interface over genre variety.

Race for the Galaxy offered economic and simultaneous-action decisions, but its corrected evaluation saturated and its games were short. A stronger opponent or a changed threshold would not satisfy the request to select a new game.

A Battle for Wesnoth campaign offered recruitment, spatial tactics, and persistent units. Its adapter would need to validate scenario transitions, scripted events, recall and gold carryover, and fog-of-war filtering. That larger scope introduced more immediate risks of interface errors and information leaks.

Fixed puzzle suites risk short search solutions or memorization. Open-ended sandboxes lacked a comparably clear native win condition. Neither was selected.

Version and references

The engine is pinned to sil-quirk/sil-q revision 4e417a7322a9cf037ee672c11f330eda7fd2c1b9, release v1.5.1-beta1 (May 16, 2026). This is a prerelease. Its release notes list bug fixes since 1.5.0 and no gameplay changes. We chose it to include those fixes, retained the upstream license and exact build identity, and tested its behavior independently. The bundled 1.5.0 manual is the manual recommended by that release.

References checked September 7, 2026:

The dated SpeedrunBench list contained SuperTux, Astray, Tuxemon, Pokemon Blue, Super Mario Land, Super Mario Bros, Wolfenstein 3D, Civilization, Mario Kart 64, and SuperTuxKart. Sil-Q was absent. This check applies to that list and date.

Plan recorded at selection

The planned sequence was:

  1. Build a protected native adapter and test scoring and information boundaries.
  2. Run GPT-5.5 at medium reasoning, review complete sessions, fix defects, and repeat validation.
  3. Play directly through the public interface.
  4. Freeze the implementation, guide, resource limits, private schedules, and analysis plan before Astra research.
  5. Complete and review ten independent Astra-medium sessions, each with eight practice games and two scored games. Report both scored outcomes and their maximum. A scored win in nine of ten sessions triggers selection of another game. Invalid or resource-bound sessions cannot establish a capability gap.
  6. If medium is valid and below that threshold, run Astra max with matched initial streams and fresh 0+2 controls. The controls compare practice plus extra inference and workspace history; they do not isolate a learning mechanism.

No model had played Sil-Q when this plan was recorded. After all ten medium sessions finished, the user reduced the follow-up to the first two launched main max attempts and canceled the controls. Both main attempts were later interrupted and excluded before scored play. The scope amendment supersedes the original follow-up plan. No replacement sessions were launched.