dataset_version stringclasses 4
values | benchmark_fingerprint stringclasses 5
values | guide stringclasses 2
values | benchmark stringclasses 1
value | map stringclasses 1
value | save stringclasses 2
values | model_id stringclasses 6
values | model_name stringclasses 6
values | endpoint stringclasses 3
values | reasoning stringclasses 2
values | max_tokens int64 32.8k 32.8k | providers listlengths 0 2 | allow_fallbacks bool 2
classes | harness_commit stringclasses 5
values | prompt_sha256 stringclasses 3
values | tools_sha256 stringclasses 2
values | game_minutes_budget float64 25 33.3 | wall_limit_minutes float64 360 360 | run_settings stringclasses 2
values | tier stringclasses 1
value | submitted_by stringclasses 1
value | series_id stringclasses 9
values | episodes int64 1 3 | episodes_completed int64 0 3 | net_worth_by_episode listlengths 1 3 | valid bool 2
classes | valid_by_episode listlengths 1 3 | score int64 3.75k 8.63k | change int64 0 1.5k | cost_usd float64 0.08 3.55 | tokens_total int64 2.5M 10.2M | game_seconds float64 1.99k 4.5k | wall_seconds float64 1.83k 6.78k | turns int64 48 204 | tool_calls int64 91 380 | playbook_final stringclasses 3
values | final_image imagewidth (px) 1.92k 1.92k | video stringclasses 9
values | files stringclasses 9
values |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1.0.0 | f1be10a5cdcc2ddf4cd395c2fcd44c862449f432d60b2019a9aff554e15f2582 | e34faf5e5efaea872f6ab4bf77ec360ee1c011d03a458702a3c2094201417c9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea-1 | z-ai/glm-5.3-flash | GLM 5.3 Flash | https://openrouter.ai/api/v1 | medium | 32,768 | [
"z-ai/fp8"
] | false | a48ab47b6d061be01042326cd332827e0bf210ec | 28202f4c82da334d8eac51d83b0ea2a661707a60041417b91a521830d97ab2a6 | a3eea35d80e92c91fb587717b28603793eead26b39125adc3685c712484a5d9c | 25 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "gameMinutes": 25, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "wallLimitMinutes": 360} | official | hesamation | 20261008T191852-c1d986e9 | 3 | 3 | [
2249,
4502,
3746
] | true | [
true,
true,
true
] | 3,746 | 1,497 | 0.383454 | 7,023,182 | 4,500.333333 | 6,782.308353 | 203 | 225 | # Oasis by the Sea-1 — playbook (final; ep1 2249, ep2 4370, ep3 3746)
## Score model (confirmed)
Net worth = gold + stored goods at sell price (wood 1, stone 7, iron 23, food 4, flour/ale 10). Selling = neutral; buying = -spread; construction = pure loss unless it produces later. Gold is king: bank goods as gold early... | v1.0/oasis-by-the-sea-construction/z-ai--glm-5.3-flash/20261008T191852-c1d986e9/episode-3/video.mp4 | v1.0/oasis-by-the-sea-construction/z-ai--glm-5.3-flash/20261008T191852-c1d986e9 | |
1.1.0 | 19121beba808ff53579a7399d99118e38af7f19074a91e7b5add9b27f1cd59f0 | e34faf5e5efaea872f6ab4bf77ec360ee1c011d03a458702a3c2094201417c9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea-1 | claude-haiku-5-5 | Claude Haiku 5.5 | https://api.anthropic.com | medium | 32,768 | [] | true | 2b90076df5f034b3a92b76343abb55c88aff578e | 6cf94b2706d5210c338bd2f6b2a1edd76ae3c5d6301b32a03b50eabba73d677a | a3eea35d80e92c91fb587717b28603793eead26b39125adc3685c712484a5d9c | 25 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "gameMinutes": 25, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "wallLimitMinutes": 360} | official | hesamation | 20261009T100849-21b4ea9d | 3 | 3 | [
4711,
5116,
5536
] | true | [
true,
true,
true
] | 5,536 | 825 | 0.30717 | 10,235,358 | 4,500.333333 | 5,126.451565 | 200 | 380 | # Oasis by the Sea: playbook (ep1 4711; ep2 5116; ep3 5536)
## Results
- Ep1: net worth about 4711. Ep2: 5116 (gold 2920 idle at end).
- Ep3 final (March 1199): gold 3760, wood 113, stone 73, iron 32, apples 104, food 104. Net worth 5536. Pop 66/66, idle peasants 14, popularity 100, tax 3.
- Gold was again the largest... | v1.1/oasis-by-the-sea-construction/claude-haiku-5-5/20261009T100849-21b4ea9d/episode-3/video.mp4 | v1.1/oasis-by-the-sea-construction/claude-haiku-5-5/20261009T100849-21b4ea9d | |
1.1.0 | e6391d4505052927c2782c445c4fbea07a2ae1c183e9bead8b5768db14da67a4 | e34faf5e5efaea872f6ab4bf77ec360ee1c011d03a458702a3c2094201417c9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea-1 | z-ai/glm-5.3-flash | GLM 5.3 Flash | https://openrouter.ai/api/v1 | medium | 32,768 | [
"z-ai/fp8",
"siliconflow/fp8"
] | false | 05376cf41ffd7aca8bfd8f581412e167723b5fab | 6cf94b2706d5210c338bd2f6b2a1edd76ae3c5d6301b32a03b50eabba73d677a | a3eea35d80e92c91fb587717b28603793eead26b39125adc3685c712484a5d9c | 25 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "gameMinutes": 25, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "wallLimitMinutes": 360} | official | hesamation | 20261009T122236-3aae6fc4 | 3 | 3 | [
4748,
4262,
5495
] | true | [
true,
true,
true
] | 5,495 | 747 | 0.514925 | 8,912,574 | 4,500.266667 | 6,370.77845 | 204 | 284 | # Oasis by the Sea — playbook (final, after ep3)
## Scores
- Ep1 ≈ 4,300 (gold 1244, iron 101, wood 217).
- Ep2 = 4,262 (gold 2093, iron 41, stone 81).
- Ep3: final reader unreadable; last verified state (2 min left): gold 828, iron 134, stone 52, food 39, pop 70, popularity 67. Estimated NW ≈ 4,600–4,800. Missed gold... | v1.1/oasis-by-the-sea-construction/z-ai--glm-5.3-flash/20261009T122236-3aae6fc4/episode-3/video.mp4 | v1.1/oasis-by-the-sea-construction/z-ai--glm-5.3-flash/20261009T122236-3aae6fc4 | |
1.5.0 | 31b5ff59f7a52c33f0dd12ebcf9bfb44649b2382825688888b82903a80868dee | 118089d290444162bb9dbb2ea535f505c0a3eb6431d399191af9e81b84a1ff9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea benchmark start | claude-haiku-5-5 | Claude Haiku 5.5 | https://api.anthropic.com | medium | 32,768 | [] | true | ee862ebfd80de51af1822f77357f5d0f4f83b2ec | 2cd353b0cf845f5d936ad2d153c5200aa21568ec3752b62a1703a79af9480ceb | 81ddda660b3a8d7c290cc8987c5f8db746da416caf60cb84ae104fd327715046 | 33.333333 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "fallbackObservation": "when-needed", "gameMinutes": 33.3333333, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "screenshotWidth": 1280, "wallLimitMinutes": 360} | official | hesamation | 20261011T141021-5ef84787 | 1 | 1 | [
5113
] | true | [
true
] | 5,113 | 0 | 0.079812 | 3,819,679 | 2,000.1 | 1,995.674441 | 76 | 136 | null | v1.5/oasis-by-the-sea-construction/claude-haiku-5-5/20261011T141021-5ef84787/episode-1/video.mp4 | v1.5/oasis-by-the-sea-construction/claude-haiku-5-5/20261011T141021-5ef84787 | |
1.5.0 | 31b5ff59f7a52c33f0dd12ebcf9bfb44649b2382825688888b82903a80868dee | 118089d290444162bb9dbb2ea535f505c0a3eb6431d399191af9e81b84a1ff9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea benchmark start | claude-sonnet-5-5 | Claude Sonnet 5.5 | https://api.anthropic.com | medium | 32,768 | [] | true | ee862ebfd80de51af1822f77357f5d0f4f83b2ec | 2cd353b0cf845f5d936ad2d153c5200aa21568ec3752b62a1703a79af9480ceb | 81ddda660b3a8d7c290cc8987c5f8db746da416caf60cb84ae104fd327715046 | 33.333333 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "fallbackObservation": "when-needed", "gameMinutes": 33.3333333, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "screenshotWidth": 1280, "wallLimitMinutes": 360} | official | hesamation | 20261011T154946-7bc1b619 | 1 | 1 | [
8633
] | true | [
true
] | 8,633 | 0 | 0.724028 | 2,498,956 | 2,000.066667 | 1,830.390383 | 48 | 91 | null | v1.5/oasis-by-the-sea-construction/claude-sonnet-5-5/20261011T154946-7bc1b619/episode-1/video.mp4 | v1.5/oasis-by-the-sea-construction/claude-sonnet-5-5/20261011T154946-7bc1b619 | |
1.5.1 | e5d9a165549e00ff5352d4bee185398365f0c48120772bc275ff4d3785bda09d | 118089d290444162bb9dbb2ea535f505c0a3eb6431d399191af9e81b84a1ff9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea benchmark start | kimi-k3 | Kimi K3 | https://api.moonshot.ai/v1 | high | 32,768 | [] | true | fe9754670075909d183d633615a137602acb47f9 | 2cd353b0cf845f5d936ad2d153c5200aa21568ec3752b62a1703a79af9480ceb | 81ddda660b3a8d7c290cc8987c5f8db746da416caf60cb84ae104fd327715046 | 33.333333 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "fallbackObservation": "when-needed", "gameMinutes": 33.3333333, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "screenshotWidth": 1280, "wallLimitMinutes": 360} | official | hesamation | 20261011T163832-bceb17f2 | 1 | 1 | [
5918
] | true | [
true
] | 5,918 | 0 | 3.54518 | 6,414,195 | 2,000.1 | 3,804.086279 | 134 | 163 | null | v1.5/oasis-by-the-sea-construction/kimi-k3/20261011T163832-bceb17f2/episode-1/video.mp4 | v1.5/oasis-by-the-sea-construction/kimi-k3/20261011T163832-bceb17f2 | |
1.5.0 | 31b5ff59f7a52c33f0dd12ebcf9bfb44649b2382825688888b82903a80868dee | 118089d290444162bb9dbb2ea535f505c0a3eb6431d399191af9e81b84a1ff9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea benchmark start | openai/gpt-6-luna | GPT-6 Luna | https://openrouter.ai/api/v1 | medium | 32,768 | [
"openai",
"azure"
] | false | ee862ebfd80de51af1822f77357f5d0f4f83b2ec | 2cd353b0cf845f5d936ad2d153c5200aa21568ec3752b62a1703a79af9480ceb | 81ddda660b3a8d7c290cc8987c5f8db746da416caf60cb84ae104fd327715046 | 33.333333 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "fallbackObservation": "when-needed", "gameMinutes": 33.3333333, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "screenshotWidth": 1280, "wallLimitMinutes": 360} | official | hesamation | 20261011T144802-e1ba2274 | 1 | 1 | [
7348
] | true | [
true
] | 7,348 | 0 | 0.113835 | 5,920,970 | 2,000.066667 | 3,170.147592 | 127 | 184 | null | v1.5/oasis-by-the-sea-construction/openai--gpt-6-luna/20261011T144802-e1ba2274/episode-1/video.mp4 | v1.5/oasis-by-the-sea-construction/openai--gpt-6-luna/20261011T144802-e1ba2274 | |
1.5.1 | e5d9a165549e00ff5352d4bee185398365f0c48120772bc275ff4d3785bda09d | 118089d290444162bb9dbb2ea535f505c0a3eb6431d399191af9e81b84a1ff9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea benchmark start | openai/gpt-6.1-sol | GPT-6.1 Sol | https://openrouter.ai/api/v1 | medium | 32,768 | [
"openai",
"azure"
] | false | fe9754670075909d183d633615a137602acb47f9 | 2cd353b0cf845f5d936ad2d153c5200aa21568ec3752b62a1703a79af9480ceb | 81ddda660b3a8d7c290cc8987c5f8db746da416caf60cb84ae104fd327715046 | 33.333333 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "fallbackObservation": "when-needed", "gameMinutes": 33.3333333, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "screenshotWidth": 1280, "wallLimitMinutes": 360} | official | hesamation | 20261011T175318-657a6989 | 1 | 1 | [
7401
] | true | [
true
] | 7,401 | 0 | 0.951924 | 4,967,660 | 2,000 | 2,705.624972 | 118 | 123 | null | v1.5/oasis-by-the-sea-construction/openai--gpt-6.1-sol/20261011T175318-657a6989/episode-1/video.mp4 | v1.5/oasis-by-the-sea-construction/openai--gpt-6.1-sol/20261011T175318-657a6989 | |
1.5.0 | 31b5ff59f7a52c33f0dd12ebcf9bfb44649b2382825688888b82903a80868dee | 118089d290444162bb9dbb2ea535f505c0a3eb6431d399191af9e81b84a1ff9f | Oasis by the Sea construction | Oasis by the Sea | Oasis by the Sea benchmark start | z-ai/glm-5.3-flash | GLM 5.3 Flash | https://openrouter.ai/api/v1 | medium | 32,768 | [
"z-ai/fp8",
"siliconflow/fp8"
] | false | ee862ebfd80de51af1822f77357f5d0f4f83b2ec | 2cd353b0cf845f5d936ad2d153c5200aa21568ec3752b62a1703a79af9480ceb | 81ddda660b3a8d7c290cc8987c5f8db746da416caf60cb84ae104fd327715046 | 33.333333 | 360 | {"contextBudget": 120000, "defaultWaitSeconds": 5, "fallbackObservation": "when-needed", "gameMinutes": 33.3333333, "imageTokenEstimate": 4096, "minTurnSeconds": 8, "recordVideo": true, "screenshotWidth": 1280, "wallLimitMinutes": 360} | official | hesamation | 20261011T132054-9bb4d8ae | 1 | 0 | [
6200
] | false | [
false
] | 6,200 | 0 | 0.145038 | 4,116,403 | 1,994.733333 | 2,625.386145 | 89 | 117 | null | v1.5/oasis-by-the-sea-construction/z-ai--glm-5.3-flash/20261011T132054-9bb4d8ae/episode-1/video.mp4 | v1.5/oasis-by-the-sea-construction/z-ai--glm-5.3-flash/20261011T132054-9bb4d8ae |
Crusader Arena runs
Runs of Crusader Arena, a benchmark in which AI models play Stronghold Crusader: Definitive Edition through screenshots and mouse and keyboard tools. Each run keeps the model's full trace, telemetry, a video and the final state of the settlement.
Research use only. Crusader Arena studies AI agents in single-player games. It is an independent project, not affiliated with or endorsed by Firefly Studios. See Responsible use and License.
The benchmark
The benchmark page
describes the task and the scoring in full; this is the short version. Every run's tables give
its exact version and settings (dataset_version, save, game_minutes_budget, run_settings).
- Task. In the scenario Oasis by the Sea construction the model starts a Free Build game on the map Oasis by the Sea from a fixed save and grows the economy. Soldiers and defences are not allowed.
- Score: net worth when the run ends: gold plus every stored good at the game's marketplace sell price. Buildings count only through what they produce. The model is told the settlement is judged on its net worth, population and popularity, without weights; the score published here is net worth.
- Time. Each episode gets a fixed amount of game time at a fixed game speed of 40, counted in game ticks. The game is paused while the model thinks, so a slow model loses no game time; a real-time limit of 6 hours only stops runs that have gone badly wrong.
- Versions 1.0 and 1.1: learning series. A model plays three episodes of 25 game minutes in
a row from the save
Oasis by the Sea-1. One thing carries over: a playbook of notes of up to 8 KB that the agent writes during play and rewrites after each episode, once it has seen its result. The last episode's net worth is the series' score; the change from episode 1 shows how much it learned. These runs keep their tables and logs only (see below). - Version 1.5: against a person's game. Single episodes with no playbook, from the save
Oasis by the Sea benchmark startfor 33 min 20 s of game time (2,000 game seconds). A person played the same save for the same game time, 25 real minutes at speed 40, and reached a net worth of 19,513 with 202 people; each episode's interview shows the model its result beside that game. - Baseline.
net_worth_growthis net worth minus what doing nothing scores on the same save and game time: 1,304 for 25 game minutes onOasis by the Sea-1and 1,220 for 33 min 20 s onOasis by the Sea benchmark start. Left alone, the starting package of 1,000 gold, 50 wood, 25 stone and 50 bread arrives and the peasants eat the bread. - The agent sees screenshots of the game window and numbers read from the running game (gold, goods, population, popularity, on-screen messages), and acts through tools to look, click, build, move the camera, trade, wait, and keep a plan and notes (see the tool reference). Each turn that runs the game lasts at least 8 game seconds. Long conversations are compacted into a handoff the model writes itself.
- Interviews (from version 1.2.0). After each episode the model is shown its result beside
a person's game on the same map and asked about its game and the benchmark, in the
conversation it played in, with the tools off.
interview.mdholds the transcript; it does not affect the score. - Versions. Benchmark versions are
MAJOR.MINOR.PATCH: a major version changes the task or scoring, a minor version what the model is told or can do, a patch fixes the harness. Runs compare within one minor version, across its patches; the accepted file fingerprints are inbenchmark-versions.json.
Load it
from datasets import load_dataset
series = load_dataset("hesamation/crusader-arena-runs", "series", split="train") # one row per series (or single episode)
episodes = load_dataset("hesamation/crusader-arena-runs", "episodes", split="train") # one row per episode
timeline = load_dataset("hesamation/crusader-arena-runs", "timeline", split="train") # one row per game minute of each episode
The tables hold the numbers and the final image. Videos, logs and model traces are files
next to them; the files and video columns give their paths in this repository.
Runs from versions 1.0 and 1.1 keep only their tables, logs and results: their videos,
screenshots (images/) and final-overview.jpg were removed, so the paths to them in
events.jsonl, manifest.json and the video column lead nowhere.
Layout
v1.0/ benchmark version, major.minor
oasis-by-the-sea-construction/ benchmark
z-ai--glm-5.3-flash/ model ID, "/" written as "--"
<series id>/ one learning series (run-<time>-<id> for a single run)
series.parquet one row: the series' score
series.json the series as the runner recorded it
playbook-after-episode-<n>.md the playbook each episode left
manifest.json every file's size and SHA-256, what was scrubbed
episode-1/
episode.parquet one row: score, game state, telemetry, final image
timeline.parquet one row per game minute: economy and telemetry so far
final-overview.jpg the settlement at the end, zoomed out
video.mp4 the run, edited, 720p: actions at real speed, idle time fast
run.json model, settings, harness version, budget, outcome
inputs.json exactly what the model was given
episode.json the scorecard
events.jsonl every model message, tool call and result
images/ the screenshots and reference images events.jsonl points to
logs.jsonl the readable log
notifications.jsonl messages the game showed
memory.json, notebook.md, playbook.md the agent's own notes
interview.md, interview.json the agent's interview after the episode
results.json, results.md the whole state of the game at the end, as data and tables
analysis.md a reading of the run written afterwards from its logs, where there is one
Comparing runs
Runs are comparable within one minor version, which is the top folder (v1.0/); patches are
harness fixes and stay comparable. dataset_version holds the exact version (1.0.0),
benchmark_fingerprint identifies the exact files a run used (a version can accept more than
one when files change without changing behaviour) and guide whether the preparation guide had
images. harness_commit, prompt_sha256 and tools_sha256 say which
code and prompt ran; run_settings holds the time budget and the other run settings.
An episode that was stopped or failed for reasons outside the model is run again; the table
holds the attempt the series kept (attempt, outcome: valid or model_failure).
tier is official for runs made by the maintainers and community for submitted runs.
A submitted score comes from the submitter's machine; the video, the screenshots in
images/ and the reader samples make it checkable, not verified.
Columns
Both tables: dataset_version, benchmark_fingerprint, guide, benchmark, map, save,
model_id, model_name, endpoint, reasoning, max_tokens, providers, allow_fallbacks,
harness_commit,
prompt_sha256, tools_sha256, game_minutes_budget, wall_limit_minutes, run_settings,
tier, submitted_by, series_id, final_image, video, files.
series: episodes, episodes_completed, net_worth_by_episode, valid (every episode
is a full-budget score), valid_by_episode, score (the last
episode's net worth), change (last minus first), cost_usd, tokens_total, game_seconds,
wall_seconds, turns, tool_calls (totals over the series), playbook_final.
episodes:
- run:
run_id,episode,attempt,outcome,episodes,started_at,ended_at,instruction,status,ended(what ended it),end_detail - score and game state:
score_source,validandinvalid(whether the net worth is a full-budget score, and if not why),net_worth,net_worth_baseline(what doing nothing scores on the same save and budget),net_worth_growth(net worth minus that baseline),goods_value,gold,goods,population,housing,popularity,total_food,structures,troops,game_year,game_month - time:
game_seconds,game_seconds_budget,wall_seconds,inference_seconds - tokens and cost:
tokens_total,tokens_input(uncached),tokens_cache_read,tokens_cache_write,tokens_output,cached_share,tokens_per_game_minute,cost_usd,cost_source(billed: what OpenRouter billed;prices: the episode's tokens at the model's list prices, for providers that do not report a cost) - agent:
turns,compactions,tool_calls,tool_calls_by_name,tool_errors,build_attempts,build_placed,build_failed,build_unverified,build_retries,build_missing,anchor_calls,anchor_placed,anchor_failed,anchor_partly_placed,peak_game_rss_mib,playbook
timeline: dataset_version, benchmark, model_id, series_id, run_id, episode,
game_minute; the reading used for that minute, sample_game_seconds (game seconds into the
budget when it was taken) and sample_source (host_observation, tool_result or
final_reading); the game then, gold, net_worth, population, housing, popularity,
total_food, structures, troops, game_year, game_month, goods; and what the agent had
spent by then, tokens_total, tokens_output, cost_usd, turns, tool_calls, tool_errors.
Row m is the latest reading at or before minute m; minute 0 is the first reading and the
last row the final one. Readings come from the agent's turns, so a minute without a fresh
reading repeats the one before (its sample_game_seconds shows this). Structures and troops
are only in full reader samples (tool_result, final_reading).
Missing values are null, never zero.
What is in the files
The logs keep everything the model was given and did: its messages and reasoning text, every
tool call and result, and the screenshots it saw. Images are files in the episode's images/
folder, and events.jsonl holds their paths: an image block reads
{"type": "image", "mimeType": "image/webp", "path": "images/<hash>.webp"}. They are the
model's images at full size, compressed to WebP at quality 70 (kept as they were when that is
no smaller), and each is stored once although the log may refer to it several times. In runs
made with the image guide (guide is not null), the preparation message also holds two
reference images: a guide to the game's construction menus, made from game screenshots, and a
screenshot of a developed settlement from an earlier human game on another map. The token
deltas of streaming replies are left out; each reply's message_end holds all of it.
Screenshots and videos show only the game window.
Before upload every text file is scrubbed: the game machine's address and paths, home
directories and host names are replaced with placeholders such as <home>; the game's process
and window IDs, the machine-wide memory and graphics details and absolute local paths are
dropped; and a bundle with any key, user name, home path, IP or email address left in it is
refused. The raw recording frames and the agent's checkpoint are not published.
Limitations
- One setup. All runs so far come from one Ubuntu laptop running the game (Steam build
- through Proton, driven by a Mac running the harness. Other machines and game builds are untested.
- Few runs. Neither the game's nor the model's randomness is fixed, so a single series mixes skill, learning and chance. Compare several series per model before drawing conclusions.
- One reading. The score is one sample of the game's memory after the final pause, read twice and checked, not an atomic snapshot.
- Submitted runs are checkable, not verified (see
tierabove). - Only one scored scenario. The other scenarios in the repository are diagnostic and have no score; military play is not scored.
See the repository's status page for what has been checked in the live game.
Responsible use
Crusader Arena is for research on AI agents in single-player games. Its game-state reader refuses multiplayer games, and the harness will not start or continue a run in one. Do not use the project against other players. The reader loads its own code into the running game, which the game's license terms may not permit; anyone reproducing these runs does so at their own risk and with their own copy of the game. The game is not included.
License
- Logs, tables, playbooks and model outputs, everything in this dataset except the game footage below: Creative Commons Attribution 4.0. Credit "Crusader Arena" and link to this dataset or the repository.
- Game footage: the screenshots and reference images in
images/,final-overview.jpgandvideo.mp4show Stronghold Crusader: Definitive Edition, © Firefly Studios. They are not covered by the license above. They are shared, unaltered apart from compression and the video's editing and overlays, for non-commercial research and to document how the agents played. Crusader Arena is not affiliated with or endorsed by Firefly Studios. Rights holders can ask for removal through the repository's issues. - The code that produced the runs is in the GitHub repository under the MIT License.
Citation
If you use these runs, please cite the repository: https://github.com/hesamsheikh/CrusaderArena.
- Downloads last month
- 193