Open-source backend for the MCPWorks platform - namespace-based function hosting and autonomous agent runtime with sandbox execution, MCP server integration, and AI orchestration
MCPWorks API presents a mixed quality profile. The server defines 12 tools with explicit schemas and descriptions, which is a strong foundation. However, there are significant gaps in schema completeness, parameter descriptions, and error handling guidance. Tool naming is generally clear and action-oriented (send_to_channel, get_state, set_state, list_state_keys, search_state, run_procedure, make_namespace, list_namespaces, make_service, list_services, delete_service, make_function). Most tools have reasonable descriptions (100-300 chars), but several critical parameters lack descriptions or have vague/incomplete constraints. The state management tools (get_state, set_state, list_state_keys, search_state) are well-designed for agents, but make_function and run_procedure have very long descriptions (500+ chars) that could confuse LLM selection. Output schemas are not explicitly documented in the tool definitions provided, which is a notable gap. Error handling descriptions are minimal, tools do not explain what to do on failure or how to recover.
Permanently delete a service and ALL its functions. This cannot be undone.
Read a single value from your persistent memory by exact key name. Use this when you know the exact key. If you don't know the key name, use list_state_keys or search_state first. Example: get_state(key='user_prefs') Returns {"key": "user_prefs", "value": ...} or an error if key not found.
List all namespaces owned by or shared with the current account. Returns namespace names and descriptions.
List all services in the current namespace. Returns service names, descriptions, and function counts. Use this to discover existing services before creating functions.
List all keys stored in your persistent memory. Use this to discover what you have saved. Takes no arguments. Returns {"keys": ["key1", "key2", ...], "count": 5, "total_size_bytes": 2048}.
delete_service lacks confirmation/dry-run pattern. A destructive operation that permanently removes all functions should require explicit user confirmation or offer a dry-run step before executing. Currently, an agent could inadvertently delete a service after a single LLM planning step.
make_function description is excessively long (800+ chars) and dense with implementation details (Python/TypeScript entry points, environment variable warnings). This violates the 10-1024 char guideline and makes LLM tool selection harder. Description should be split: a brief 50-100 char summary, with detailed entry-point info in a separate documentation field or parameter descriptions.
Output schemas are not documented for any of the 12 tools. LLMs cannot plan downstream tool calls or extract the right data without knowing what fields to expect. For example, make_service, make_function, and list_services should all document their return structures (success status, resource IDs, metadata).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2026-07-28+ | v2 |
Create a new function in an existing service. The service must already exist (use make_service first). Workflow: 1) make_service if needed → 2) make_function with service name → 3) execute via run server. The 'service' parameter is just the service name (e.g. 'utils'), NOT a namespace or fully-qualified path. Python entry points (in priority order): 1) 'result = ...' — assign to result variable. 2) 'output = ...' — alias for result. 3) 'def main(input):' — function receiving input dict, return value is the result. 4) 'def handler(input, context):' — function receiving input dict and context dict. TypeScript entry points: 1) 'export default function main(input) { ... }' — default export (preferred). 2) 'export default function handler(input, context) { ... }' — with context. 3) 'module.exports.main = function(input) { ... }' — CommonJS. 4) 'const result = ...' — simple assignment. NEVER hardcode API keys, tokens, secrets, or credentials in code — use required_env for caller-provided secrets or agent state (via context['state']) for stored secrets.
Create a new namespace for organizing services and functions. A namespace is the top-level container — you must have one before creating services or functions. Your current namespace is already set by the MCP server connection URL.
Create a new service within the current namespace. A service is a group of related functions. You must create a service before creating functions. The service is automatically created in the namespace this MCP server is connected to — do NOT pass a namespace parameter. Example: make_service(name='utils') → make_function(service='utils', name='hello', ...)
Execute a procedure step-by-step with enforced function calls. The orchestrator steps through each defined step, calling the required function and capturing its result as proof before advancing. ALWAYS use this instead of calling functions directly when a matching procedure exists. Example: run_procedure(service='social', name='post-bluesky-thread', input_context={'post1_text': '...', 'post2_text': '...'}) Returns {"success": true, "steps_completed": 3, "final_text": "..."}.
Search your persistent memory by keyword. Finds keys and values containing the search term (case-insensitive). Use this when you need to find something but don't remember the exact key name. Example: search_state(query='project') Returns {"matches": [{"key": "my_project", "preview": "...first 100 chars..."}], "query": "project", "total_searched": 10}.
Send a message to a configured communication channel. Use this to notify users or post updates. Example: send_to_channel(channel_type='discord', message='Deploy complete!') Returns {"sent": true, "channel": "discord"} on success.
Save a value to your persistent memory. Values survive restarts and are available in future conversations. The value can be any JSON type: string, number, boolean, object, or array. Example: set_state(key='last_run', value='2026-03-20T10:00:00Z') Returns {"key": "last_run", "stored": true} on success. Special keys: '__soul__' (your identity), '__goals__' (your objectives), '__heartbeat_instructions__' (what to do on next heartbeat wake).
run_procedure has an undefined concept of 'procedure' in its description. LLMs cannot know what a procedure is, what services/names are valid, or what input_context structure is expected without external documentation or discovery examples. The description should clarify: 'A procedure is a named sequence of steps defined in the specified service. Steps are executed sequentially; each step calls a function and passes its result to the next step.'
Error handling is absent or minimal across all tools. For example: get_state does not explain what to do if a key is not found (should it suggest search_state? list_state_keys?), make_service does not explain what happens if the namespace doesn't exist, and delete_service has no recovery guidance. Error responses should categorize failures (retryable, user-fixable, fatal) and suggest next steps.
Parameters lack type constraints and validation hints. For example: 'message' in send_to_channel has no max length or format hint; 'query' in search_state has no max length; 'name' in make_service has no pattern constraint (unlike make_namespace which specifies a regex). LLM descriptions should include explicit constraints: '(max 500 chars)', '(lowercase, alphanumeric, hyphens)', '(max 100 chars)'.
Example values are embedded in descriptions (e.g., send_to_channel example, set_state example with '2026-03-20T10:00:00Z'). LLMs tend to reuse example values literally in real calls, causing failures. Replace examples with enum constraints, format hints, or move examples to a separate examples field.
make_function schema lacks 'required' array, making it unclear which parameters are actually required vs optional. Comparing parameter descriptions with schema: 'service', 'name', and 'backend' appear required, but this should be explicit in the schema's 'required' field.
search_state and list_namespaces do not support pagination. If an agent has many stored keys or namespaces, results could exceed context limits. Add optional 'limit' and 'offset' or 'cursor' parameters, and document in the description.
Parameter descriptions are sometimes vague or incomplete. For example: run_procedure's 'input_context' is described as 'Optional initial context available to step 1' without clarifying its structure or format. make_function's 'config' parameter says 'Backend-specific configuration. Not needed for code_sandbox' but doesn't define what fields are valid for other backends.