One MCP for developers - No tool tax, no context rot. 100+ tools including Brave, Google, Context7, Excalidraw, Version Checker, Excel, File Ops, Database, Playwright, Chrome DevTools and many more.
OneTool MCP has severe definition quality issues. The 4 tools provided have inconsistent schema completeness, vague descriptions, and unclear parameter documentation. The 'run' tool is particularly problematic, it accepts arbitrary code/function calls with minimal validation guidance, creating a foot-gun for agent usage. The 'security' tool's description is confusing (is it for checking if code is safe, or for learning security rules?). The 'server' tool has 4 optional parameters with minimal documentation of their mutual exclusivity or effects. The 'skills' tool has better structure but still lacks clarity on what 'skill' means in this context. No tool provides output schema documentation. Error handling and recovery guidance is absent across all tools. The codebase suggests 100+ additional bundled tools exist but are NOT defined in the submitted schema, this makes evaluation impossible for most of the toolkit. Conservative scoring applied due to missing context and vague, sometimes contradictory descriptions.
Speak `text` aloud using the macOS `say` command.
Check security rules for code validation.
List or inspect runtime proxy server state.
Load a single image into session storage and return a stable handle.
Tool 'run' has dangerously vague description and parameters. Description says 'Execute code and tool calls' but does NOT clarify which code is safe, what happens on syntax errors, whether side effects are reversible, or how the allowlist-based security interacts with arbitrary Python code execution. The 'command' parameter lacks type constraints, format examples, or validation rules.
Missing output schemas for ALL tools. The rubric requires 'Document the output schema. LLMs need to know what fields to expect.' None of the 4 tools define what they return, no fields, types, or structure. This forces LLMs to guess at response structure and breaks downstream tool chaining.
Parameter descriptions are missing or generic. Tool 'server' has 4 parameters ('status', 'enable', 'disable', 'restart') with NO descriptions. LLMs cannot determine which parameter to use or how the mutual exclusivity works. Tool 'security' parameter 'check' says 'Pattern to check (e.g. 'os', 'json.loads', 'pickle.*')', this is an example, not a real description of what 'pattern' means or what happens when it's empty.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 17 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Tool 'run' with IRREVERSIBLE risk has no error handling or recovery guidance. If code execution fails, the description provides no hint on how to recover, what error types to expect, or whether to retry. The error classification pattern is completely absent.
Tool 'run' name violates verb_noun naming convention (rubric: 'Start tool names with a verb'). 'run' is too generic, it does not distinguish between 'run Python code', 'run a function call', or 'run a shell command'. A 4-character name provides zero context. Rename to 'execute_code' or 'execute_function_call'.
Tool 'security' description is self-contradictory. It says 'Check security rules for code validation' (suggesting it validates code) but then says 'If empty, returns summary of all security rules' (suggesting it's a discovery tool). Is this a validator or an info tool? LLMs will be confused.
Tool 'skills' parameter 'info' accepts 'list', 'min', 'full' but this is documented as an example inside the description text, not as an enum constraint. LLMs may hallucinate other values like 'short', 'detailed', 'verbose'.
Bundled tools are NOT schema-registered. Codebase mentions '100+ tools' but none appear in the submitted tool list. This makes evaluation of the actual toolkit impossible. Tool discovery and composition patterns cannot be assessed.