AI Models are benchmarked on their ability to play Starcraft: Brood Wars
From the member
GBTI NetworkBrood War Bench is an experiment by software engineer Ben Swerdlow that tests whether modern AI agents can independently play StarCraft: Brood War, Blizzard’s 1998 real-time strategy expansion to StarCraft. Instead of controlling the game directly with a mouse and keyboard, the models issue commands through an agent interface and must continuously manage resources, build armies, research technology, scout, and fight while the game keeps running.¹ ²
Swerdlow tested models from OpenAI, Anthropic, and xAI and found that none played beyond a beginner level. Codex Astra performed best in the benchmark, while Claude Fable tended to build more complete economies and technology trees. Grok frequently spent too long reasoning between actions, sometimes failing to build or attack at all. One of the most interesting findings was that some agents effectively treated the real-time game like a turn-based one, continuing to think while their opponents were actively attacking them.
The benchmark is published through swerdlow.dev, Swerdlow’s personal site for software, AI, infrastructure, and experimental projects. Brood War makes an unusually useful AI test because success requires more than answering correctly. An agent has to observe, decide, act, recover from mistakes, and keep doing all of that under continuous time pressure.³
Footnotes
- Ben Swerdlow, Brood War Bench. (Agent StarCraft)
- Blizzard Entertainment, StarCraft: Brood War, a 1998 expansion adding new units, campaigns, maps, and other features to StarCraft. (classic.battle.net)
- Ben Swerdlow, swerdlow.dev.

0 Comments
No comments yet. Be the first. Members comment from the GBTI local client, where comments are submitted as pull requests and auto-published for paid members.
Become a memberComments are for members. Become a member to join the conversation, or if you already are.
You are signed in as a member. .