MCP server for MITRE Caldera — async, typed, universal
caldera-mcp demonstrates strong naming conventions (verb-first, action-oriented) and comprehensive parameter schemas with descriptions. Most tools have clear, actionable descriptions (averaging 120-160 chars). Schemas are well-structured with proper JSON Schema types. However, there are gaps in output documentation, limited error guidance, and no tool annotations. The server uses fastmcp correctly with proper typing via Pydantic models, but tool definitions rely on indirect registration patterns that limit verifiability from static code. Composition is sound, each tool has single responsibility, and tool chains are viable (e.g., list_abilities → create_adversary → create_operation). No secrets in parameters. Missing: per-tool return type documentation, error recovery guides, and input validation examples in descriptions.
Manually add a link (ability execution) to a running operation.
Create a new adversary profile. An adversary is an ordered list of abilities that form an attack scenario. The order of ``ability_ids`` defines execution sequence.
Create a new operation. An operation executes an adversary profile against a group of agents. By default it starts in ``paused`` state.
Find adversaries by name or by ability ID. At least one filter must be provided. Searches are case-insensitive for name, exact match for ability_id.
Get full details of a specific ability.
Get full details of a specific adversary profile.
Missing output schema documentation for all tools. Tool definitions do not explicitly document the structure of returned objects (fields, types, nesting). LLMs cannot reliably plan downstream operations without knowing what fields a tool returns.
Minimal error guidance in descriptions. Tools like get_ability, get_adversary, get_agent, get_operation do not document what happens when the resource is not found or how the LLM should recover. Descriptions lack 'try X() instead' recovery hints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 51 | - | v1 |
Get full details of a specific agent by its PAW identifier.
Get the execution result of a specific link in an operation. The output is automatically base64-decoded.
Get details of a specific operation.
Check whether the Caldera API is reachable. Returns a dict with ``healthy`` (bool) and ``url``.
List Caldera abilities with optional filters. Filters are applied dynamically — no hardcoded values. Pass any tactic, technique_id, or plugin that exists in your Caldera instance.
List all Caldera adversary profiles.
List all Caldera agents.
List all Caldera operations.
List all available payloads on the Caldera server.
Free-text search across all abilities. Searches ability name, description, technique name, and tactic. Case-insensitive. Returns up to ``limit`` results.
Return Caldera MCP server information. Includes version, configured URL, available tactics (discovered dynamically from loaded abilities), agent count, and planner list.
Change the state of an existing operation.
Generic descriptions for list/discovery tools. list_adversaries, list_operations, list_payloads have descriptions under 70 characters with no context on when to call them or what structure they reveal.
No tool annotations present. Tools like create_adversary, create_operation, update_operation_state, add_link_to_operation are write/destructive but lack destructiveHint annotation. read_only tools (list_*, get_*, search_*) lack readOnlyHint. Per current MCP spec, annotations guide agent planning.
Indirect tool registration obscures schema verifiability. Tools are registered via module-level register() functions (e.g., register_abilities, register_adversaries in server.py) rather than explicit inline definitions. Static analysis cannot fully verify all parameter schemas and descriptions without executing or reading each registration function.
Pagination not addressed in list tools. list_abilities, list_adversaries, list_agents, list_operations, list_payloads have no limit, offset, or page parameters.
No input validation guidance in parameter descriptions. Parameters like adversary_id, ability_id, paw, operation_id lack format hints. Descriptions should specify format (UUID, alphanumeric), length, and constraints so LLMs avoid invalid inputs.