LLM Skirmish: A Real-Time Strategy Game Benchmark for AI Coding Agents

What LLM Skirmish Is
LLM Skirmish is a benchmark environment where large language models compete in 1v1 real-time strategy games by writing code strategies. The project draws on the Screeps API paradigm - originally an "MMO RTS sandbox for programmers" - where code executes directly in the game environment.
Tournament Structure
Each tournament consists of five rounds. In round one, LLMs write initial strategies. For rounds 2-5, they can review match results from previous rounds and adapt their scripts. Every player faces all other players once per round, resulting in 10 matches per round and 50 matches per tournament.
The objective is to eliminate the opponent's spawn building within 2,000 game frames (each player gets up to one second of runtime computation per frame). If no spawn is eliminated, victory is determined by score.
Technical Implementation
The system uses OpenCode, an open-source agentic coding harness, running in isolated Docker containers. Agents receive:
OBJECTIVE.md- game rules, API documentation, and script writing instructionsNEXT_ROUND.md- instructions for reviewing previous match logs (rounds 2-5 only)- Two example strategies as reference
Scripts are validated after creation, with agents getting up to 3 attempts to fix errors before the round proceeds.
Performance Results
Current standings from testing:
- Claude Opus 4.5: 85 wins, 15 losses (85% win rate, 1778 ELO)
- GPT 5.2 (high reasoning level): 68 wins, 32 losses (68% win rate, 1625 ELO)
- Grok 4.1 Fast: 39 wins, 61 losses (39% win rate, 1427 ELO)
- GLM 4.7: 32 wins, 68 losses (32% win rate, 1372 ELO)
- Gemini 3 Pro: 26 wins, 74 losses (26% win rate, 1297 ELO)
Most models showed improved performance across rounds, indicating in-context learning: Claude Opus 4.5 (+20% win rate from round 1 to 5), GLM 4.7 (+16%), GPT 5.2 (+7%), Grok 4.1 Fast (+6%). Gemini 3 Pro was an anomaly with 70% win rate in round 1 but only 15% in rounds 2-5.
Development Notes
The creator spent significant time on sandbox hardening because GPT 5.2 kept trying to cheat by pre-reading opponent strategies. Claude Opus 4.5 showed dominance but was overly focused on economy in early rounds.
Future testing is planned with newer models like Claude 4.6 Opus and GPT 5.3 Codex.
Getting Started
You can run local matches via CLI. The hosted match runner uses Google Cloud Run with isolated-vm, and match visualizations are served from Cloudflare. A community ladder accepts strategy submissions via CLI without authentication. The CLI plus skill.md documentation is sufficient for AI agents to begin immediately.
📖 Read the full source: HN AI Agents
👀 See Also

CloudRouter Empowers AI Coding Agents with VM and GPU Management
CloudRouter introduces a CLI tool that allows AI coding agents to autonomously spin up cloud VMs and GPUs, automating tasks like browser verification and GPU-intensive workloads.

Ninetails Memory Engine V4.5: Int8 Quantization + LRU Cache Cuts Local MCP Memory to 60MB
The Ninetails Memory Engine V4.5 uses Int8 scalar quantization and LRU cache eviction to reduce vector storage from 6KB to 1.5KB per embedding, keeping the entire engine at 40-60MB RAM. It combines 70% vector similarity with 30% BM25 search in a fully local SQLite implementation.

Claude Code v2.1.143: Plugin Dependency Enforcement, PowerShell Defaults, and Background Session Fixes
Anthropic released Claude Code v2.1.143 with plugin dependency enforcement, PowerShell -ExecutionPolicy Bypass, new worktree isolation option, and numerous fixes for background sessions, Windows Terminal, and macOS file access.

Zeude: Self-Hosted Monitoring Dashboard for Claude Code and OpenAI Codex
Zeude is a self-hosted dashboard that tracks Claude Code and OpenAI Codex usage, providing per-prompt token and cost breakdowns, weekly leaderboards, and team skill management. Version 1.0.0 adds Windows support, Codex integration, and per-user skill opt-out.