Web UI Agent Platform based on Claude Code - an MCP server that enables Claude to interact with local file systems, execute bash commands, manage todos, perform web searches, and coordinate complex multi-step tasks
This MCP server exposes 15 tools with significant definition quality gaps. Most tools lack complete parameter descriptions, and many have schemas that are either missing or inferred from scattered React components rather than centralized tool registration. Tool naming follows some conventions (Read, Edit, Write, Bash) but lacks consistency (exit_plan_mode vs ExitPlanMode are aliases with the same description). Critical security concerns: the Bash tool accepts arbitrary commands with no validation, whitelist, or execution warnings. File operation tools (Read, Edit, MultiEdit, Write) lack clear permission/safety boundaries. The server bundles unrelated concerns: file editing, web search, task execution, and todo management in a single tool set. Schema inspection reveals incomplete definitions, several tools have only partial input schemas visible in the code. Parameter descriptions exist but are often minimal (10-30 chars), falling below the 72-char baseline for A+ tools. Output schemas are not documented for any tool. Error handling is not visible in the tool definitions themselves.
Execute a bash command and return the output
Edit a file by replacing old_string with new_string
Exit planning mode and execute the plan (alias for exit_plan_mode)
Find files matching a glob pattern
Search for a pattern in files using grep
List files and directories in a path
Apply multiple edits to a single file
Read file content with optional offset and limit parameters
Bash tool accepts arbitrary shell commands with no validation, sandboxing, whitelist, or safety warnings. LLMs can be tricked into executing dangerous commands (rm -rf /, privilege escalation, data exfiltration).
Duplicate tools with same functionality: exit_plan_mode and ExitPlanMode both accept a 'plan' parameter and exit planning mode. LLMs waste reasoning cycles deciding between them. No explanation for why two aliases exist.
Missing output schemas for all 15 tools. LLMs cannot predict what fields to expect or plan downstream tool chaining. No documentation of return types visible in source.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 27 | 1.17.0+ | v1 |
Execute a complex multi-step task with child messages and nested tool use
Read todo items
Write or update todo items
Fetch content from a URL
Perform a web search
Write content to a new file
Exit planning mode and execute the plan
Parameter descriptions are consistently too short (10-40 chars vs. 72-char baseline). E.g. 'The path to list' for LS, 'The bash command to execute' for Bash. Missing context on format, constraints, range, or examples.
File operation tools (Read, Edit, MultiEdit, Write) lack any permission checks or safety warnings. No documentation of path traversal protection, symlink handling, or access control. LLMs could be induced to read/write sensitive files.
TodoRead has empty input schema ({}). No parameters means the tool always returns the same static list, unclear when/why to call it, and no flexibility to filter or paginate.
Edit, MultiEdit, Write, and Bash are marked WRITE/IRREVERSIBLE but have no confirmation step, dry-run mode, or undo capability. No error guidance on failure (e.g. file not found, permission denied, edit not found).
Tool definitions are scattered across React components (ToolContent.tsx, SearchTool.tsx, WebTool.tsx) rather than centralized MCP server registration. Schemas appear inferred rather than explicitly registered with the MCP SDK. Tool name/description/schema mapping is implicit.
Grep and Glob tools require both 'pattern' and 'path' parameters but no documentation on glob syntax vs. regex syntax, or whether 'path' is a directory or file glob prefix. LLMs may pass invalid patterns.
WebSearch and WebFetch take only a 'query'/'url' parameter, no rate limiting, timeout, result limit, or documentation of retry behavior. LLMs could trigger runaway searches or long-hanging requests.
Task tool accepts only a free-text 'description', no structured input, no clear success/failure output, no sub-task breakdown visible. Highly ambiguous.