MCP Server with FastAPI backend providing filesystem operations, shell command execution, and LLM code generation tools with authentication, audit logging, and monitoring
This server defines 7 tools via fastmcp with basic parameter schemas present. However, multiple critical gaps significantly undermine quality: (1) Tool names lack clarity, 'file_system_create_directory_tool' is verbose and doesn't follow verb_noun pattern; names should be 'create_directory', not 'file_system_create_directory_tool'. (2) Descriptions exist but are extremely terse (10-30 chars), well below the 50-200 char production baseline and insufficient for LLM decision-making. Examples: 'Create a directory within the sandbox' (38 chars) and 'Write content to a file within the sandbox' (42 chars) lack context on when/why to use vs. alternatives, or what the tool returns. (3) Parameters lack descriptions in the visible source, 'path' and 'content' are named but have no type specs or constraints visible in the tool definitions shown. (4) No output schema documented for any tool, LLMs cannot plan downstream calls or extract results. (5) No error handling strategy visible, what happens on permission denied, disk full, or invalid paths? (6) Shell execution tool is high-risk but lacks granular permission checks, dry-run support, or command validation guidance. (7) LLM tools (OpenAI/Gemini) lack API key management documentation and rate-limit guidance. The server demonstrates competent infrastructure (Docker, type checking, tests) but tool definitions themselves are below production standards.
Execute a shell command within the sandbox with security restrictions.
Create a directory within the sandbox.
List contents of a directory within the sandbox.
Read content from a file within the sandbox.
Write content to a file within the sandbox.
Generate code using Google Gemini model with specified programming language and requirements.
Generate code using OpenAI's GPT model with specified programming language and requirements.
Tool names are overly verbose and inconsistent with verb_noun pattern. Examples: 'file_system_create_directory_tool' should be 'create_directory'; 'execute_shell_command_tool' should be 'execute_command'. Production baseline: 90% of A+ tools use action-verb prefixes with avg length 18 chars. These average 35+ chars and contain redundant scope prefixes.
Descriptions are critically short (10-42 chars) and lack LLM decision context. Production baseline: 194 chars avg (p10=34, p90=392). Current descriptions do not answer WHAT the tool does, WHEN to use it vs. alternatives, or WHAT it returns. Example: 'Create a directory within the sandbox' leaves LLMs unable to decide when to call this vs. write_file or other filesystem ops.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions are missing or minimal in source code. Visible schema shows {'path': {'type': 'string', 'description': '...'}} but descriptions are terse (10-20 chars). Production baseline: 72 chars avg per param. Example: 'path' params lack guidance on absolute vs. relative, symlink handling, or path traversal protections.
No output schemas documented for any tool. LLMs cannot plan downstream calls or know what fields to extract. Production requirement: 100% of A+ tools have documented return types. Unknown: Do these tools return JSON with structured fields, plain text, or status codes? This forces agents to guess.
Shell command execution tool (execute_shell_command_tool) lacks error handling, dry-run, or confirmation flow. High-risk irreversible operations (rm, kill, etc.) should support confirmation patterns. No visible recovery guidance, if a command fails, what should the agent do next?
LLM code generation tools (OpenAI/Gemini variants) do not document API key/secret injection, rate limits, or cost implications. Credentials must never be exposed as parameters per pattern:secret-injection. No visible indication of how API keys are managed.
Filesystem tools lack validation constraints. No enum for file permissions, no min/max lengths for paths, no regex patterns. Example: 'path' parameter should document max length (e.g. 255 chars), allowed characters, and symlink handling. Unbounded/undocumented params invite invalid input.
No error classification or recovery guidance visible. When a tool fails (file not found, permission denied, timeout), the agent has no way to know: is this retryable? User-fixable? Fatal? Production pattern: categorize errors and include actionable next steps.
Tool composition risk: no output schema means downstream tools cannot safely chain. Example: does 'list_directory' return file names as a list? An object with metadata? Unknown return structure blocks agent planning for multi-step operations like 'list files, then read the first one'.