A robust, extensible server for managing, versioning, and serving prompts and templates for LLM applications, built on the Model Context Protocol (MCP)
This evaluation is based on the provided tool definitions from the mcp-devtools-unified server. The server declares 7 tools but critical quality gaps severely limit production readiness. Tool naming follows verb_noun convention but descriptions are minimal (13-80 chars, well below the 194-char baseline). No input parameter descriptions are provided, parameters exist but lack context for LLM usage. Output schemas are completely undocumented. Error handling guidance is absent. The server appears to be a development/analysis toolset but lacks the polish required for reliable agent integration. Specific tool examination reveals consistent pattern: naming is adequate, but descriptions are bare and parameter documentation is entirely missing.
Analyze code changes using git diff + static analysis
Perform unified static code analysis on source files
Start a debugging session with GDB/LLDB
Execute analysis tools in Docker containers
Analyze git commit history for debugging insights
Execute tests with automatic framework detection
Execute a multi-tool orchestrated workflow
Parameter descriptions completely missing across all tools. Every parameter (files, language, config, since_commit, tools, directory, framework, pattern, executable, debugger, args, breakpoints, file, author, since, until, image, command, volumes, working_dir, workflow_name, context) lacks a description explaining what it controls. LLMs cannot infer parameter meaning from names alone, required by pattern:tool-description.
No output schemas documented for any tool. LLMs need to know what fields to expect in responses to plan downstream calls and extract relevant data. analyze_code, analyze_changed_code, run_tests, debug_session_start, git_analyze_history, docker_exec_analysis, and workflow_execute provide no return type documentation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Tool descriptions are below actionable length. 'Perform unified static code analysis on source files' (56 chars), 'Analyze code changes using git diff + static analysis' (54 chars), 'Execute tests with automatic framework detection' (48 chars), 'Start a debugging session with GDB/LLDB' (38 chars), 'Analyze git commit history for debugging insights' (48 chars), 'Execute analysis tools in Docker containers' (41 chars), and 'Execute a multi-tool orchestrated workflow' (40 chars) all fall below the 194-char baseline and lack WHEN to use it, prerequisites, or return value guidance.
No error handling or recovery guidance. Tools performing analysis, testing, and debugging (debug_session_start, docker_exec_analysis, workflow_execute) can fail in multiple ways but provide no indication of what to do next. LLMs need to know: is it retryable? Should I ask the user? Is it fatal?
Destructive/write operations (debug_session_start, docker_exec_analysis, workflow_execute) lack confirmation/dry-run patterns. An LLM could start a debugging session or execute arbitrary Docker commands without a chance to verify intent. Irreversible operations should support a dry-run or confirmation step per pattern:confirmation-request.
Ambiguous parameter 'config' in analyze_code. This is a generic object with no description, no type hints for nested fields, and no explanation of what configuration options are valid. Undocumented parameter dependencies cause silent misuse per pattern:tool-description.
Parameter 'framework' in run_tests lacks enum constraint. Free-form string with no list of valid frameworks (pytest, unittest, jest, junit, etc.) invites hallucinated values. Should declare as enum per pattern:constrained-input.
Parameter 'debugger' in debug_session_start lacks enum constraint. Should declare valid debuggers (gdb, lldb) as an enum instead of free-form string.
Tool naming ambiguity: 'analyze_code' and 'analyze_changed_code' are similar but their distinction is unclear from names alone. The difference (changed vs all code) is critical but undocumented. LLMs conflate similar names per pattern:tool-chain.
No parameter validation rules documented. How many files can analyze_code accept? What is the max size? Is there a language whitelist? LLMs need explicit constraints to avoid invalid input per pattern:tool-description.