A modular code execution framework for AI agents with sandboxed execution, adapters for external APIs, and pre-execution code analysis
MCX server exhibits significant quality gaps across naming, descriptions, and schema completeness. Of 25 tools examined, most lack comprehensive parameter descriptions and have minimal output schema documentation. Naming is mostly action-verb based (positive), but descriptions are sparse and underly the 10-1024 character baseline. No evidence of error handling guidance, input validation details, or permission declarations. The template tools (listItems, getItem, createItem, updateItem, deleteItem) appear to be scaffold examples rather than production implementations. Chrome DevTools tools (launchChrome, connect, killChrome) lack parameter validation documentation and error recovery paths. Supabase adapter tools have basic descriptions but omit critical details about defaults, pagination, and output formats. Overall definition quality is below the median baseline of 45-55.
Apply a database migration. Auto-selects current project.
Connect to Chrome DevTools Protocol
Create a new item
Create a new project (auto-selected for subsequent calls)
Delete an item
Execute SQL query. Auto-selects current project.
Generate TypeScript types from schema. Auto-selects current project.
Get item by ID
Sparse tool descriptions: 16 of 25 tools have descriptions under 50 characters or trivial content (e.g., 'Connect to Chrome DevTools Protocol' with no context on preconditions, return value, or error paths).
Missing output schema documentation: No tool in the provided code explicitly documents its return type, field structure, or pagination format. LLMs cannot infer what fields to expect from responses, forcing them to guess and leading to incorrect downstream tool chaining. Critical for patterns like get_organization → select_project.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 30 | - | v1 |
Get security or performance advisors. Auto-selects current project.
Get project API keys. Auto-selects current project.
Get service logs. Auto-selects current project.
Get organization details
Get project details. Auto-selects current project if not provided.
Get project API URL. Auto-selects current project.
Kill Chrome process
Launch Chrome with remote debugging enabled
List all items with optional filtering
List database migrations. Auto-selects current project.
List all organizations
List all projects
List tables in schemas. Auto-selects current project.
Pause a project. Auto-selects current project.
Restore a paused project. Auto-selects current project.
Select a project for subsequent commands (from list_projects)
Update an existing item
No parameter constraints or validation guidance in descriptions. Parameters like 'region' in create_project, 'service' in get_logs, and 'type' in get_advisors lack enum/regex/range documentation. LLMs will guess invalid values (e.g., 'us-west-5' for region when only 'us-east-1, us-west-1' are valid).
Error handling is not documented. No tool description explains what errors are retryable, what to do if a resource is not found, or how to recover from common failures. For example, execute_sql could fail due to schema mismatch, query syntax, or permission; no guidance provided.
Destructive operations lack confirmation or dry-run patterns. Tools like 'deleteItem', 'pause_project', 'killChrome', and 'apply_migration' can cause data loss or service disruption but provide no warning mechanism or rollback guidance for agents.
No permission or scope declarations on any tool. Tools exposing API keys (get_api_keys), executing SQL (execute_sql), and killing processes (killChrome) should declare required permissions (read:secrets, write:db, admin:process). Enables least-privilege agent configs.
Template tools (listItems, getItem, createItem, updateItem, deleteItem) appear to be placeholder examples from adapter.template.ts, not production implementations. No evidence these are wired into actual backend services. If these are meant to be examples only, they should be removed or clearly marked as scaffolds.
Vague parameter names and missing type hints. 'get_logs' accepts a 'service' parameter described as 'Service: api, postgres, edge, auth, storage, realtime', but is this an enum or a free-form string? Should be declared as enum in schema. Similarly, 'type' in get_advisors is ambiguous: does it accept strings or IDs?
Pagination guidance missing. Tools returning lists (list_organizations, list_projects, list_tables, list_migrations, listItems) do not document whether pagination is supported, how to specify limit/offset, or what total count looks like.
Chrome DevTools integration tools lack stateful session context documentation. launchChrome, connect, and killChrome suggest session management via 'currentTargetId' and 'session' variables, but this statefulness is not documented in tool descriptions. Agents cannot know that connect requires launchChrome to be called first or that state persists across calls.