An MCP client manager that orchestrates multiple MCP servers, retrieves tools via hybrid search with reranking, generates YAML workflow plans using an LLM, validates plans against available tools, and executes them. Includes a filesystem server component.
This MCP server exposes 11 filesystem tools with clear action-verb naming and moderate descriptions. All tools have explicit input schemas with types and required fields. However, descriptions are terse (averaging 80-120 chars), parameter descriptions lack specificity about constraints and ranges, output schemas are not documented, and error handling lacks recovery guidance. The tools follow basic Agentic Tool Patterns but miss production-grade polish in parameter validation rules, error classification, and output schema documentation. No tool combines multiple concerns (composition is good), but parameter descriptions need expansion to guide LLM behavior around path restrictions, encoding, and error modes.
Create a new directory. Creates parent directories as needed. Only works within allowed directories.
Recursively list all files and directories in a tree structure. Only works within allowed directories.
Efficiently edit a file by specifying multiple old/new text replacements. Supports dry_run to preview changes before committing. Only works within allowed directories.
Get metadata about a file (size, timestamps, permissions). Only works within allowed directories.
List all files and subdirectories in a directory. Only works within allowed directories.
Move or rename a file or directory. Only works within allowed directories.
Parameter descriptions lack constraint documentation. The 'path' parameter appears in all tools but descriptions never specify: allowed directory restrictions, path format requirements (absolute vs relative), encoding assumptions, or what happens if path escapes allowed dirs. LLMs cannot infer these constraints and will pass invalid paths.
No output schema documentation for any tool. Descriptions state what tools do but never document what fields the LLM should expect in responses. For example, list_directory returns file/dir items but response structure is invisible. Without documented schemas, LLMs cannot reliably extract or chain outputs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Read the complete contents of a file asynchronously. Supports UTF-8 encoding and raises detailed errors if the file cannot be read. Only works within allowed directories.
Read the contents of multiple files asynchronously. Returns each file's content prefixed with its path, separated by '---'. Continues on individual file errors. Only works within allowed directories.
Search for files matching a pattern within a directory. Uses case-insensitive matching. Only works within allowed directories.
Dynamically update the list of allowed directories for filesystem access. Only works within allowed directories.
Write content to a file. Creates the file if it does not exist, or overwrites it if it does. Only works within allowed directories.
Error handling lacks recovery guidance. Descriptions mention errors ('raises detailed errors if the file cannot be read') but do not specify what types of errors occur, how to categorize them (retryable vs user-fixable), or what the LLM should do next. No examples like 'If path_traversal_error, verify the path is within allowed directories.'
set_allowed_directories lacks security context. Description says 'Dynamically update the list of allowed directories' but does not specify: Is there a permission gate? Can any caller update allowed dirs? What happens if an LLM expands allowed_directories to '/'? This is a sensitive tool and needs explicit permission/validation guidance.
edit_file dry_run parameter lacks clarity on use case. Description mentions 'supports dry_run to preview changes before committing' but does not explain when an LLM should call with dry_run=true vs committing directly, or how the agent should interpret dry_run output. Best practice: 'Call with dry_run=true to preview edits without modifying the file. Returns the modified content. Once verified, call again with dry_run=false to commit.'
search_files pattern and exclude_patterns parameters lack specificity. Descriptions say 'case-insensitive matching' and 'Regex patterns to exclude' but do not document: What regex dialect (Python re, PCRE, etc.)? Does pattern match full path or filename? How should LLMs construct regex patterns? This forces LLMs to guess or will cause unexpected results.
read_multiple_files continues on error but description does not document partial-failure response format. If 3 of 5 files read successfully, what does the response contain? All successes? Mixed success/failure objects? Without documented structure, LLM cannot reliably handle partial results.
No pagination or result-limit guidance for list_directory and directory_tree. If a directory contains 10,000 files, does list_directory return all 10,000? Should the LLM expect this to blow the context window? No mention of limits, pagination, or truncation behavior.