The server defines 4 tools with JSON Schema input validation, but quality is uneven. Tool names are clear and action-oriented (generate_*, ask_*), but descriptions are brief and lack LLM-optimization guidance. Parameter descriptions exist but are minimal. Output schemas are not documented anywhere in the provided source. Error handling is present (ErrorCode, McpError imports) but the actual error recovery guidance in tool implementations is not visible in the snippet provided. The example file shows tool invocation via curl, suggesting HTTP transport, but the actual MCP server implementation details in src/index.ts are truncated mid-function.
Tool descriptions are under 50 characters and lack actionable context. 'Generate code based on a description' and 'Ask a question to the LLM' provide minimal guidance on WHEN to use each tool or how it differs from others. Descriptions should be 50-200 chars and answer: what does it do, when should the LLM call it, and what does it return?
Output schemas are not documented. The tools return LLamaIndex ChatResponse/MessageContent objects (inferred from imports), but the JSON structure and fields returned to the MCP client are not visible in the provided source. LLMs cannot plan downstream calls without knowing what fields to expect.
Expand each tool description to 80-150 characters, following this template: '[ACTION] by [METHOD]. Use this when [WHEN]. Returns [WHAT STRUCTURE]. Common parameters: [KEY PARAMS].' Example: 'Generate code via LLM. Use this to produce new functions, classes, or scripts. Returns generated_code (string) and language_detected (string). Requires: description (what to build); Optional: language, additionalContext (constraints).'
Add enum constraints to 'language' parameter: {"type": "string", "enum": ["JavaScript", "Python", "TypeScript", "Java", "Go"], "description": "Programming language for code generation. Supported: JavaScript, Python, TypeScript, Java, Go."}, include in both description text and schema.
Document output schemas for each tool. Example for generate_code: {"type": "object", "properties": {"generated_code": {"type": "string", "description": "The generated code"}, "language": {"type": "string", "description": "Detected or specified language"}, "explanation": {"type": "string", "description": "Brief explanation of what the code does"}}, "required": ["generated_code"]}
Expand parameter descriptions with format/range constraints. Example for 'description' in generate_code: 'Description of the code to generate (1-500 characters). Be specific: "function to calculate factorial" vs generic "write a function". Avoid example values like "hello world", instead describe the intent.'
Implement path sanitization for generate_code_to_file. Validate filePath is within a designated output directory (e.g. /tmp/generated/) and reject '..' or absolute paths outside that scope. Return a clear error: 'Cannot write to [path], output files must be within /tmp/generated/ directory.'
Score history
Overall score trend
↑ 23 points across a rubric change (v1 → v2)
54/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
54
2026-07-28+
v2
2026-03-09
F
31
-
v1
Parameter descriptions are minimal (1-10 words). 'Description of the code to generate', 'Programming language', 'Additional context or requirements' lack specificity. For 'language', should include: 'Supported languages: JavaScript, Python, TypeScript, etc.' For 'additionalContext', should specify length limits, format, and when to include it.
generate_code_to_file accepts 'filePath' (a user-provided filesystem path) without visible sanitization or validation. This is a path-traversal risk. The tool should validate the path is within an allowed directory, or use a resource URI abstraction instead.
No enum constraints on 'language' or 'format' parameters. LLMs will hallucinate unsupported language names (e.g. 'Kotlin', 'R', 'Pascal') if only a few are actually supported by the LLM provider. Define enums to restrict to supported values.
No visible error recovery guidance. The code imports McpError and ErrorCode but the actual error responses from each tool (e.g., when LLM request fails, file write fails, language unsupported) are not shown. Error messages should tell LLMs what to do next: 'Unsupported language. Supported: JavaScript, Python, TypeScript. Call ask_question to suggest an alternative.' instead of a bare error code.
Add per-tool error handling examples in the server code. When LLM call fails, return: {"error": "LLM_UNAVAILABLE", "message": "LLM provider returned status 503. Retry in 30s or try a different model via LLM_MODEL_NAME."} instead of a generic 500 error.
Implement confirmation step or dry-run for generate_code_to_file. This tool modifies the filesystem, add a 'dry_run' boolean parameter that previews the change without writing, or wrap the write in a confirmation prompt that shows the diff before commit.
Add 'idempotentHint' annotation to read-only tools (generate_code, generate_documentation, ask_question) in the tool registration if using tool annotations (currently not enabled but recommend adding). This signals to agents these calls are safe to retry.
Document LLM configuration dependencies. The server relies on environment variables (LLM_MODEL_NAME, LLM_MODEL_PROVIDER, etc.) but the tool descriptions do not mention fallback behavior if these are not set. Add a note: 'Requires LLM_MODEL_PROVIDER (ollama, openai, or bedrock) and LLM_MODEL_NAME environment variables to be configured.'