Self-extending MCP server - build and execute custom AI tools at runtime
Skillz provides 14 tools with explicit schemas and descriptions. Most tools have clear action verbs (build_tool, register_script, delete_tool, call_tool, create_pipeline, import_tool). However, several tools exhibit significant definition quality issues: (1) Descriptions are often terse and lack context for when to use the tool instead of alternatives. For example, 'List all available tools' (list_tools) is only 26 chars and provides no guidance on when an agent should call it. (2) Parameter descriptions are present but inconsistent in depth, some tools like register_script have extensive guidance (e.g., 'CRITICAL for Python: Use sys.stdin.readline() NOT sys.stdin.read()') while others like explore_skill lack prerequisite context. (3) Output schemas are not documented in the source, only input schemas are visible. (4) Error handling guidance is absent, tools return results or failures, but descriptions do not explain recovery paths. (5) Some high-risk tools (build_tool, execute_code, delete_tool) lack clear confirmation/dry-run patterns despite their destructive potential. The tool naming is strong (verb-leading, clear intent), and input parameters are well-constrained with enums and type definitions where applicable. Baseline analysis: 14/14 tools have descriptions (100%, exceeds 94% A+ baseline); all have JSON Schema input definitions; but output schemas and error guidance are absent. Tool names average ~15 chars and start with action verbs, meeting naming baselines.
Unified tool building/management for WASM and script-based tools
Call a tool with arguments
Create a pipeline that executes multiple tools in sequence
Delete a value from tool memory storage
Delete a tool from the registry
Code execution mode - compose multiple tools via code
Explore a skill's resources and capabilities
Get a value from tool memory storage
Output schemas not documented. Tool definitions specify input_schema but no output_schema is visible in registration or descriptions. LLMs cannot predict what fields to expect from tool responses, forcing trial-and-error interpretation and wasting tokens parsing unstructured output.
Destructive tools lack confirmation/dry-run patterns. delete_tool, delete_memory, and execute_code (with arbitrary code execution) have no confirmation step, dry-run mode, or explicit 'this will modify state' language in descriptions. Agents can trigger irreversible changes without safeguards.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | 2024-11-05+ | v1 |
Import a tool from an external source
List all available tools
Register an external MCP server
Register a script tool
Set a value in tool memory storage
Manage tool versions - list, rollback, or get info about versions
Error handling guidance absent. Tool descriptions do not explain what to do if a call fails. For example, 'Call a tool with arguments' (call_tool) provides no guidance on retryability, required preconditions, or recovery paths when the tool execution fails.
Inconsistent parameter guidance depth. register_script includes critical guidance ('CRITICAL for Python: Use sys.stdin.readline() NOT sys.stdin.read()') while explore_skill has minimal context. LLMs may misuse tools when guidance is sparse, especially for interpret-level decisions (e.g., which language interpreter to invoke).
Terse descriptions on discovery tools. list_tools (26 chars) and explore_skill (50 chars) lack context on when to call them. Baselines show A+ tool descriptions average 194 chars; these fall well short (18% and 26% of baseline), limiting LLM's ability to decide when discovery is necessary.
Missing context on tool chaining. create_pipeline returns a pipeline object, but the response structure is not documented. Agents cannot predict whether the response includes a pipeline_id, status, or other fields needed for downstream calls (e.g., call_tool). Broken chaining forces discovery detours.
Parameter dependencies undocumented. tool_version's 'version' parameter is only required when action='rollback', but the description does not state this. LLMs may omit the parameter for 'list' or 'info' actions and waste a round-trip on error clarification.