Toolkit for building, recording, and sharing mock APIs with API simulation, contract testing, digital twin capabilities, and MCP integration
The server provides 11 mock API management tools with reasonable structure. All tools have names starting with action verbs (list_, start_, stop_, get_, create_, update_, delete_, run_, generate_) and descriptions present. However, descriptions are generic and brief (10-50 chars typically), parameter descriptions are minimal, output schemas are not documented in the source, and error handling guidance is absent. The tools follow a consistent naming pattern for the domain, but lack the depth of documentation expected for production agent usage. STDIO-only transport is a hard limitation, this server cannot be used by hosted MCP clients and cannot be remotely tested.
Creates a new mock service from a YAML definition.
Deletes a mock service.
Generates a mock service from an AI prompt.
Retrieves the latest request logs for a service.
Retrieves the OpenAPI specification for a service.
Gets the current status of a mock service.
Lists all available mock services.
Runs contract-driven API tests against a service.
Output schemas not documented. No evidence in source code of return type specifications for any of the 11 tools. LLMs cannot plan downstream operations or validate results without knowing the structure of responses (e.g., does list_services return an array of service objects, or strings, or complex metadata?). This violates the critical check: 'Document the output schema.'
Descriptions are too generic and brief (typically 25-50 chars). Examples: 'Lists all available mock services' (32 chars), 'Starts a specific mock service' (31 chars). These do not answer WHEN to use each tool, what prerequisites exist, or how they interact. Baseline for A+ tools is 50-200 chars with explicit context. Current descriptions lack depth to guide LLM selection among similar tools (e.g., get_service_status vs get_service_logs both retrieve service data but descriptions don't explain the distinction).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Starts a specific mock service.
Stops a running mock service.
Updates an existing mock service configuration.
Parameter descriptions are minimal. 'service_name' appears in 8 tools but is described identically as 'The name of the service to X'. No guidance on format (alphanumeric, length limits, allowed characters), whether it's case-sensitive, or how to discover valid names. The rubric requires: 'Describe the expected format, range, and allowed values directly in the parameter description.' Current descriptions violate this.
No error handling guidance documented. No indication of what errors each tool might return, when to retry, or recovery steps. The rubric requires: 'Error responses must tell the LLM what to do next' and 'Categorize errors as retryable, user-fixable, or fatal.' The source shows no error classification or recovery hints for tools like delete_service (destructive) or create_service (which requires valid YAML).
Destructive operations (delete_service) lack confirmation/dry-run patterns. The rubric states: 'Irreversible operations (delete, send, publish) should support a dry-run or confirmation step.' No evidence of a confirm_delete_service or dry_run parameter, and no error guidance if a service is accidentally deleted.
YAML parameter (create_service, update_service) has no validation rules documented. LLMs will attempt to construct YAML from natural language, likely generating invalid or malformed definitions. Rubric: 'Describe the expected format, range, and allowed values directly in the parameter description.' Should include: expected YAML structure, required fields, example skeleton, or link to schema.
No pagination or result limiting guidance. get_service_logs accepts 'limit' parameter (good) but the description does not state: default value, min/max bounds, or what happens if limit is omitted or set to 0. The rubric requires: 'Specify minimum and maximum for numeric parameters (e.g. page_size 1 - 100).' Also, other list tools (list_services) show no pagination parameters despite domain suggesting potentially large result sets.
No tool composition guidance. No evidence that output from one tool chains cleanly to inputs of another. For example: does list_services return service_name in a format that start_service accepts directly? If create_service returns an ID, do update_service and delete_service accept it? Missing documentation forces LLMs to infer or attempt discovery calls.
generate_service_from_prompt lacks context on what constitutes a valid prompt and what model/framework generates the YAML. Is this calling an LLM internally? What are output guarantees? Current description 'Generates a mock service from an AI prompt' is circular and gives no actionable guidance.