Mock Fleet MCP demonstrates solid foundational quality with 32 well-named, documented tools operating on WireMock and Kubernetes Fleet APIs. All tools follow consistent verb_noun naming conventions (list_*, get_*, create_*, update_*, delete_*, start_*, stop_*) and include descriptions. Input schemas are present with types and descriptions for all parameters. However, several gaps prevent a higher score: (1) Output schemas are not documented in the provided source, response structures are inferred but not formally specified, making downstream tool composition harder for LLMs. (2) Parameter descriptions, while present, are often terse (10-30 chars) and lack actionable constraints, e.g., 'limit' is described as 'Page size' without stating the valid range (1-100?). (3) Error handling details are absent from the definitions; no guidance on retryability, recovery steps, or what happens on concurrency conflicts (resourceVersion mismatches). (4) Tool composition could be stronger, tools like import_mock_configs require understanding of Kubernetes resourceVersion patterns, which is not documented for LLMs unfamiliar with optimistic concurrency control. (5) No per-tool risk annotations in the source code (only inferred from tool names and input params).
Tools (32)
count_requestsread onlysource verified75/100
Count logged requests in a running mock pod matching request criteria (WireMock tool).
create_stubwritesource verified75/100
Create a new stub in a running mock pod (WireMock tool).
delete_body_filedestructivesource verified77/100
Delete a binary body file from a running mock pod (WireMock tool).
Output schemas not documented in tool definitions. Response structures for all 32 tools are inferred from API behavior but not formally specified. This breaks downstream tool composition, LLMs cannot plan multi-step workflows when they don't know what fields a tool returns.
Parameter descriptions are terse and lack actionable constraints. 'limit' is described as 'Page size' without specifying valid range (1 - 100? 1 - 1000?). 'cursor' is 'Opaque continuation cursor', too vague. Constraints like min/max, allowed enum values, and format requirements must be in the description text for LLM consumption (they cannot parse JSON Schema fields).
Recommendations
Document the response schema for every tool. For list_* tools, specify the structure of returned items (e.g., list_mocks returns [{mockId: string, status: 'running'|'stopped', config: {...}}]). For get_* tools, document all returned fields. This is essential for LLM planning.
Expand parameter descriptions to include constraints and examples (without making them prescriptive). Change 'Page size' to 'Page size (limit results returned; valid range 1 - 100, default 20)'. Change 'Opaque continuation cursor' to 'Opaque continuation cursor (pass the value from a previous response's next_cursor to fetch the next page)'.
Add recovery guidance to tool descriptions for all write/destructive operations. Example for delete_mock_config: 'Atomically delete a saved mock configuration using optimistic concurrency control. If you receive a conflict error (409), fetch the latest resourceVersion via get_mock_config and retry. This tool is DESTRUCTIVE, deleted configurations cannot be recovered.'
Document the apply strategy enum: 'Apply strategy must be "futureOnly" (default, applies only to future pods) or "restartActive" (restart running mocks immediately). Use "futureOnly" for non-urgent changes, "restartActive" to force immediate effect.'
For import_mock_configs, add a usage example in the description: 'Atomically create or replace saved mock overrides from a configuration array. First call export_mock_configs to fetch the current resourceVersion, then pass it here along with modified mocks to apply.'
Score history
Overall score trend
First recorded score · v2 rubric
68/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
C
68
2025-06-18+
v2
get_body_fileread onlysource verified77/100
Get a binary body file from a running mock pod (WireMock tool).
get_mock_configread onlysource verified78/100
Get one mock's configuration and Fleet routing metadata.
get_near_missesread onlysource verified73/100
Get near-miss request/stub pairings from a running mock pod (WireMock tool).
Get recording status in a running mock pod (WireMock tool).
get_stubread onlysource verified82/100
Get one stub's configuration from a running mock pod (WireMock tool).
import_mock_configswritesource verified73/100
Atomically create or replace saved mock overrides from a configuration array using optimistic concurrency. Preserves other mocks and applies to future pods only.
list_body_filesread onlysource verified77/100
List binary body files stored in a running mock pod (WireMock tool).
list_mocksread onlysource verified80/100
List configured and active Mock Fleet mocks without starting any mock pod.
Optimistic concurrency control (resourceVersion) is undocumented for LLM use. Tools update_mock_config, delete_mock_config, and import_mock_configs require a 'resourceVersion' parameter from get_mock_config responses, but this Kubernetes pattern is not explained. LLMs will not understand why these calls fail with 409 Conflict or how to recover. Needs explicit recovery guidance: 'If you get a concurrency conflict, call get_mock_config to fetch the latest resourceVersion and retry.'
No documented error handling or recovery paths. Tool descriptions do not explain failure modes (e.g., 'Mock not found', 'WireMock pod unreachable', 'ConfigMap conflict') or what LLMs should do next (retry? lookup? ask user?). Errors returned by the server must include actionable next steps.
Tool composition gaps. Import/export tools require understanding of array format and resourceVersion. Stub operations require 'mockId' and 'stubId', but there is no guidance on obtaining a stubId if only a request path is known. Tools should explicitly document lookup sequences: 'To update a stub by path, call list_stubs first to find its UUID.'
Create/update distinction not clarified. update_mock_config creates or updates based on whether the mock exists, but this overloading is not documented. Parameter 'baseline' is 'optional for update', when is it required? When should LLMs provide it? Adds ambiguity.
Risk annotations present in input but not in source code. Mock ID and stubId are string types with no validation hints. LLMs cannot verify format (UUID? kebab-case? arbitrary string?) before calling tools. Descriptions should state: 'Mock ID: lowercase alphanumeric and hyphens, 1-63 chars' or similar.
No documented prerequisites or state dependencies. send_request, for example, silently fails if the mock pod is not running, but the description does not warn LLMs. Similarly, snapshot_requests requires that recording was active. Tools should state: 'Requires mock to be running (call start_mock first if needed).'
Add idempotency guarantees: document which tools are safe to retry without side effects (e.g., get_* and list_* are read-only; create/update/delete are idempotent if resourceVersion matches). Tools without explicit idempotency guarantees invite duplicate side effects on retries.
Document what 'matching criteria' means for find_requests, count_requests, get_near_misses. Is it WireMock MatcherDef JSON? A query DSL? LLMs cannot infer this from the name alone.
Add prerequisites to stub and request tools: 'Requires the mock to be in the running state. If the mock is stopped, call start_mock first. If not found, call list_mocks to verify the mock exists.'
For baseline and user parameters in update_mock_config, specify the accepted object structure: 'Baseline configuration (required for create): {...WireMock baseline config...}' with a reference to the WireMock config schema or a canonical example.
Document snapshot_requests more explicitly: 'Save recorded requests as JSON stubs. This tool converts candidate stubs from a prior stop_recording call into persistent stubs. Requires that recording was active and has been stopped; use get_recording_status to verify readiness.'
Add error classification: for each tool, document retryable errors (e.g., 'Conflict: resourceVersion mismatch, retry after fetching latest') vs user-fixable errors (e.g., 'Mock not found, verify mockId via list_mocks') vs fatal errors.
For put_body_file, clarify base64 encoding: 'Binary body content (base64 encoded). Decode on the server side; do not pass raw binary data.'
Add a discovery pattern to the descriptions: tools like list_option_definitions and list_mocks should state their purpose in tool discovery: 'Call this first to discover available mock configurations before creating or updating.'
Document the relationship between start_recording, stop_recording, and snapshot_requests: 'Recording workflow: start_recording -> send requests -> stop_recording -> snapshot_requests to save. Recording is not persisted until snapshotted.'
For all tools accepting mockId, add a hint: 'If you don't have the mock ID, call list_mocks to search by name or pattern.'
Document page size defaults and limits. E.g., 'limit: Page size (default 20, max 100). Larger limits are capped by the server to prevent context overflow.'
Add timeout guidance: tools that wait for pod operations (start_mock, stop_mock) should document timeout behavior: 'Times out after 30 seconds if the pod does not transition to the target state. Returns a timeout error; safe to retry.'
For reset_request_journal and reset_scenarios, add a confirmation warning: 'DESTRUCTIVE: Permanently clears all recorded requests/scenarios. This cannot be undone. Ensure you have exported data if needed.'