Model Context Protocol Local Development - A sandbox for managing local development environments
Four tools are explicitly registered in server.py with names, descriptions, and input schemas. However, the evaluation reveals systematic gaps in description depth, parameter documentation, and output schema clarity. All four tools have descriptions but they are minimal (26-61 chars), falling short of the 50-200 char LLM-optimized baseline. Input schemas are present and typed, but parameter descriptions are extremely terse (7-28 chars). Most critically, output schemas are not documented, the tool implementation returns JSON but the description does not specify the output structure, forcing LLMs to infer what fields are returned. Error handling exists (e.g., 'Unknown environment' messages) but is minimal and does not guide recovery. The server appears functional but underoptimized for agent reasoning.
Clean up a local development environment
Create a new local development environment from a filesystem path
Create a new local development environment from a GitHub repository
Auto-detect and run tests in a local development environment
Minimal parameter descriptions. All parameters have terse descriptions (7-28 characters). Parameter descriptions must explain format, range, and allowed values. Example: 'github_url' is described only as 'GitHub repository URL', it should specify expected format (https://github.com/owner/repo), required protocol, and what constitutes a valid URL.
Output schemas are not documented. Tool descriptions and source code show JSON responses (success, data, coverage fields) but the tool descriptions do not specify the output schema. LLMs cannot plan downstream tool calls or extract data without knowing the response structure. Example: local_dev_from_github returns {success, data:{id, working_dir, created_at, runtime}}, this must be declared in the tool definition.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Tool descriptions are too brief (26-61 characters, baseline is 50-200 LLM-optimized). Descriptions lack context for when and why to use each tool, and do not state preconditions or side effects. Example: 'Clean up a local development environment' does not explain that cleanup is destructive, what happens to the environment, or how to verify it succeeded.
No risk annotations. Destructive operations (local_dev_cleanup) and write operations (local_dev_from_github, local_dev_from_filesystem) lack tool annotations to declare their risk profile. Current spec supports destructiveHint, idempotentHint, and readOnlyHint, these should be used to signal to agents which tools modify state and which are safe to retry.
Sparse error guidance. Error responses like 'Unknown environment: {env_id}' are informative but do not guide recovery. LLMs receive the error but have no hint about next steps (e.g., 'List available environments with list_environments', if such a tool existed). Error messages should include recovery guidance.
No idempotency declaration. Tools that create or modify environments (local_dev_from_github, local_dev_from_filesystem) should declare whether repeated calls with identical arguments produce the same result or trigger duplicate creation. If non-idempotent, agents cannot safely retry on transient failure.