Static source inference · medium confidence · detected: stateful session
Deprecated protocol patterns detected
Summary
The server defines 9 tools with explicit schemas and descriptions visible in src/mcp.rs. Tool naming follows verb_noun convention (file_read, file_write, editor_search, shell_exec). Descriptions are present but brief (10-60 chars), falling below the 50-200 char production baseline. Input schemas are properly structured with type definitions and required fields. However, descriptions lack context about when to use each tool, prerequisites, and what to do on failure. Parameter descriptions are minimal ('File path', 'Directory path'). No output schemas are documented. Error handling is present but basic, tools return string errors without recovery guidance. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk levels (READ_ONLY, WRITE, IRREVERSIBLE). Security-sensitive tools (shell_exec, shell_spawn, file_write) lack confirmation/dry-run patterns. Overall, the server is functional but below production baseline for LLM integration.
Tools (9)
editor_replacewritesource verified65/100
Replace a unique substring in a file. Fails if not found or if multiple matches (ambiguous)
editor_replace_lineswritesource verified62/100
Replace a range of lines [start, end] inclusive. 1-indexed
editor_searchread onlysource verified70/100
Search for text in a file, returns matching lines with line numbers
file_listread onlysource verified68/100
List directory contents. Directories have trailing /
file_readread onlysource verified68/100
Read a file and return its contents as a string
file_writewritesource verified67/100
Create or overwrite a file. Parent directory must exist
Descriptions are too brief (10-60 chars, baseline 50-200) and lack context. None answer: What does this do? When to use it instead of a similar tool? What are prerequisites? LLMs cannot infer optimal tool selection from bare descriptions like 'Execute a shell command and return its output'.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk levels documented in the metadata (READ_ONLY, WRITE, IRREVERSIBLE). This forces LLMs to infer destructiveness from text alone, increasing hallucination risk.
Expand all tool descriptions to 50-200 characters, answering: (1) What does this tool do in one sentence? (2) When would you use it instead of similar tools (e.g., editor_replace vs editor_replace_lines)? (3) What does it return? Example: 'file_read: Read and return the complete text contents of a file on disk. Use when you need to inspect a file's current state before editing. Returns the full file content as a string.'
Add detailed parameter descriptions for all inputs. Specify format (e.g., 'must be an absolute path or ~/... relative to sandbox root'), constraints (e.g., 'must be valid UTF-8'), and implications. Example for editor_replace: 'old_text: The exact substring to find. Search is case-sensitive. Must be unique, if multiple matches exist, the operation fails with an ambiguity error.'
Add tool annotations to each tool definition: readOnlyHint=true for file_read, file_list, editor_search, shell_exec (when non-modifying); destructiveHint=true for file_write, editor_replace, editor_replace_lines, shell_kill; idempotentHint=true for file_read, file_list, editor_search.
Implement a confirmation pattern for destructive operations. Add a 'dry_run' or 'confirm' parameter (boolean, default=false) to file_write, editor_replace, editor_replace_lines, shell_exec, shell_kill. When true, return what would happen without executing. Example: 'Would replace 5 lines in file.txt (lines 10-14)'.
Document all output schemas. For each tool, specify the response object structure. Examples: file_read returns {'content': string}; file_list returns {'entries': [{name: string, is_dir: bool}], 'total': int}; editor_search returns {'matches': [{line_number: int, line_text: string}]}.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Stateful initialize / Mcp-Session-Id (removed; protocol is stateless) - make each request self-contained
Score history
Overall score trend
↑ 18 points across a rubric change (v1 → v2)
50/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
50
<=2025-11-25
v2
2026-03-09
F
32
-
v1
irreversiblesource verified58/100
Kill a process by PID
shell_spawnirreversiblesource verified60/100
Spawn a background process (detached, no output capture)
Irreversible tools (shell_exec, shell_spawn, shell_kill, editor_replace, editor_replace_lines) lack dry-run or confirmation patterns. Agents can and will invoke destructive operations without safeguards, leading to data loss.
Error responses return raw string messages ('cannot read {path}: {e}') without recovery guidance. LLMs cannot determine if an error is retryable, requires user input, or is fatal. No guidance like 'Try search_files() first'.
No documented output schemas for any tool. Callers (LLMs) do not know what fields to expect in responses, forcing them to guess field names and types. This breaks composition, if file_read returns 'content' vs 'file_content', chaining fails.
Parameter descriptions are absent or one-word generic ('File path', 'Directory path', 'Text to search for'). Production baseline requires descriptions to include format, constraints, and examples. LLMs cannot infer valid inputs from bare names.
shell_exec and shell_spawn accept arbitrary shell commands as free-form strings with no validation. LLMs can be tricked into injecting command sequences. No mention of safety considerations, escaping, or injection risks in descriptions.
editor_replace_lines parameter 'start' and 'end' are documented as 'number' type but descriptions say '1-indexed, inclusive'. No clarification of edge cases (what if start > end? what if end exceeds file length?). LLMs will pass invalid ranges.
editor_replace_lines
Enhance error messages with recovery guidance. Replace 'cannot read {path}: {e}' with 'Cannot read {path}: {specific_error}. Verify the path exists and is readable. Try file_list() first to browse available files.'
Add validation and constraints to numeric parameters. For editor_replace_lines: validate start >= 1, end >= start, end <= file_line_count. Return detailed errors: 'Invalid range: start=10, end=5. Start must be <= end.'
Document path handling explicitly. Clarify: What characters are allowed? How are symlinks handled? Does the tool resolve ~/user/docs correctly? Are relative paths allowed? Example: 'path: Absolute or sandbox-relative path. Supports ~/ prefix for user home. Paths are resolved relative to the sandbox root. No access outside sandbox.'
Add rate limiting and timeout guidance. For shell_exec and shell_spawn, document: 'Commands timeout after 60 seconds. Max output is 64 KB. Long-running operations should use shell_spawn (detached) to avoid blocking.'
Split shell_exec and shell_spawn into explicit separate tools with clearer semantics. Currently, the distinction between 'wait for output' vs 'detached' is unclear. Add tool: 'shell_exec_with_timeout(command, timeout_ms, max_output_bytes)' and 'shell_spawn_background(command) -> pid'.