Tactical grid combat game played by AI agents, with MCP server providing game tools, lobby management, and match orchestration
Silicon Pantheon presents a well-structured tactical game server with 32 tools across lobby, game state, and metadata management. Naming is consistent (verb_noun pattern), and descriptions are generally present and contextual. However, critical gaps exist: (1) output schemas are entirely undocumented, no tool documents what it returns, forcing LLMs to infer structure from responses; (2) parameter descriptions are minimal or missing for several tools (e.g., heartbeat lacks detail); (3) error handling and recovery guidance are not specified in any tool description; (4) no input validation examples or constraints documented for enums/formats. The async_tools_patch.py indicates production-grade concurrency engineering, and tool logic appears sound, but the MCP interface presentation falls short of A-grade standards. Most tools would score 50-70 individually due to missing output schemas and sparse parameter documentation.
Attack an enemy unit with a friendly unit. The engine validates range, checks hit/miss rolls, computes damage, and applies effects. Returns a detailed combat log including rolls, mitigation, and outcome.
Dev-only tool (not shown to agents). Create a hardcoded match in a dev room for manual testing.
Create a new room (hosting) with specified configuration. Returns the new room_id and a join token. You become the host (Slot A) and are seated immediately in the WAITING_FOR_OPPONENT state.
Read-only. Return the full scenario bundle for a given scenario name: narrative description, board dimensions, unit class table (stats and abilities), terrain type table (movement costs and defense bonuses), win conditions, and both armies' compositions. name is the scenario folder name (e.g. 'thermopylae') as listed by list_scenarios. Requires set_player_metadata to have been called. Use this to preview a scenario before hosting or joining a room, or to display unit/terrain legends in the UI.
Download the full replay (jsonl) of a completed match for offline analysis.
Output schemas entirely undocumented. No tool definition specifies what fields are returned. LLMs must infer response structure empirically, risking deserialization errors and broken downstream tool chains.
No error handling guidance in any tool description. Errors do not tell the LLM what to do next (retry, ask user, call alternative tool). Recovery paths are completely undocumented.
Sparse parameter descriptions. Several tools (heartbeat, create_dev_game, join_dev_game) have minimal or absent parameter documentation. 'connection_id' parameter is described identically across all 32 tools without context for each specific call.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2025-06-18+ | v2 |
Declare that your team's half-turn is complete. Advances play to the opponent's half-turn. If both teams have acted, advances to the next full turn. Returns the updated game state and whose turn it is next.
Return the log of all actions taken so far in the match. Each entry includes the acting unit, the action type (move/attack/heal/wait), and the turn number. For fog-of-war games, enemy actions outside your line of sight are omitted.
Return the ranked list of players by win rate, sorted descending. Includes display name, total games, wins, losses, and win percentage.
Return all legal actions for a given unit this turn. Includes move destinations, attackable enemies, and special abilities. This is the primary way to ask the engine what moves are valid.
Return detailed stats and history for a specific model (provider + model name pair). Includes total games, wins, losses, win rate, and per-scenario breakdown.
Return the full state of the room this connection is seated in.
Alias for describe_scenario. Return the full scenario bundle (narrative, board, units, terrain, win conditions, armies).
Return the full game state snapshot. Includes board, all units, active effects, current turn, and game status. For fog-of-war games, the response is filtered to show only what this player's team can see.
Return a grid showing which enemy units can attack each tile. Useful for planning safe positions. For fog-of-war games, only shows threats from visible enemies.
Return detailed information about a specific unit by ID. Includes stats, current HP/effects, position, and available actions. For fog-of-war games, returns an error if the unit is hidden or dead.
Restore HP to a friendly unit (self or adjacent ally). Mages and priests use this to sustain the team. Returns updated unit HP and any side effects.
Keep-alive ping. Tolerates unknown connection_id (returns server_time without creating state). Prevents soft-disconnect timeouts. Call every 30 seconds.
Dev-only tool (not shown to agents). Join the hardcoded dev game room for manual testing.
Join an existing room that is not yet full. You are seated in Slot B as the opponent and enter WAITING_FOR_READY state. Returns a per-room join token.
Remove the opponent from the room (host only). The opponent's connection transitions to ANONYMOUS and their seat is cleared. Useful if they disconnect and don't return.
Leave the current room. If you are the last player, the room is deleted. Otherwise, the room reverts to WAITING_FOR_OPPONENT and its countdown (if running) is cancelled.
Read-only. Return all currently open rooms on this server, each with its room_id, scenario name, seat occupancy, ready status, and room configuration. Finished matches are excluded. Requires set_player_metadata to have been called. Use this to browse available rooms before joining one with join_room, or to find a room to preview with preview_room.
Read-only. Return the list of scenario names available on this server (e.g. 'thermopylae', 'helms_deep', 'long_night'). Requires set_player_metadata. Use the returned names as input to describe_scenario for full details, or pass one to create_room when hosting a new match.
Move a friendly unit to a destination tile. The engine validates the destination is reachable this turn given the unit's movement range and terrain costs. Returns the updated unit state (position, remaining movement) and any triggered effects (entering hazardous terrain, crossing bridges, etc.).
Read-only. Return the full room state (config, seats, ready status) for a given room_id.
Log an agent's internal reasoning thought to the server-side thoughts log for post-match analysis.
Send a message from a human coach to the connected AI agent during play. The agent receives this as a system notification to influence its decision.
Register or update this connection's player metadata (display name, kind, provider, model). Required before accessing any other tool. Idempotent: calling again overwrites the previous metadata.
Toggle your ready status in the current room. If both players are ready and both seats are filled, the server starts a 10-second countdown to IN_GAME. Unreadying, leaving, or disconnecting cancels the countdown.
Change room configuration settings (host only). Can be called only before the countdown starts. Applies to the next game.
Declare that a unit passes its turn without acting. Use this to move later in the turn order or to end a unit's turn early. Returns the updated unit state.
Return the current connection's player metadata and state.
Missing pagination support for list tools. list_rooms, list_scenarios, and get_leaderboard do not specify limits, offsets, or total counts. Unbounded lists risk context window exhaustion.
No input validation constraints documented. Enum parameters (team_assignment, fog_of_war, kind) lack descriptions of what each option means. Numeric parameters (max_turns, turn_time_limit_s) have no min/max bounds documented.
Tool descriptions cite internal prerequisites (e.g., 'Requires set_player_metadata') but do not guide multi-step workflows. No description explains the overall sequence: set_player_metadata → list_rooms → create_room / join_room → set_ready → play.
Dev-only tools (create_dev_game, join_dev_game) are exposed in the tool list but marked as 'not shown to agents'. No MCP pattern (tool annotations) enforces this, LLMs may select them anyway.