Unified tool surface with AI, cloud control, vector search, and multi-framework UI
The Hanzo MCP server has 4 tools with definitions visible in the source code. However, critical quality issues significantly impact its score. While input schemas are present with documented types and enums for most tools (agent, browser, code, config), the descriptions are problematic: they are extremely verbose (1500+ characters for 'agent' and 'browser'), violate the 10-1024 character guideline, and bury actionable intent under excessive detail. The 'agent' tool description reads like a feature dump rather than a focused explanation of when to use it. Parameter descriptions are present but inconsistently detailed. Error handling guidance is absent, tools provide no recovery hints, retryability classification, or actionable error messages. The tool design shows composition issues: 'agent' combines multiple orchestration patterns (run, dag, swarm, consensus, dispatch, zen, review) into a single tool rather than splitting by concern. Output schemas are not documented anywhere in the provided source. The code tool has good schema coverage but the config tool mixes multiple configuration planes (index.scope, tools.*.enabled) into one tool. No security patterns are evident (no audit trails, permission gates, or scope declarations). Overall, while schemas exist, the descriptions are antithetical to LLM usability, and critical patterns (error guidance, output schemas, composition) are missing.
Multi-agent orchestration tool (HIP-0300): spawns native coding CLIs (claude/codex/gemini/grok/qwen/dev) and Anthropic-compatible providers as tokio subprocesses, then composes them: run (single agent), dag (dependency graph), swarm (fanned template), consensus (multi-model agreement), dispatch (parallel agents), list/status/config (introspection), zen (64-path oracle guidance), review (constructive review by focus area)
Browser automation tool (HIP-0300): Full Playwright surface driven by one long-lived Node driver process. Handles ~90 actions: navigation (navigate, reload, go_back, go_forward, close), content operations (content, url, title, set_content), input actions (click, dblclick, type, fill, clear, press, select_option, check, uncheck, upload), mouse/touch interactions (hover, drag, mouse_move, scroll, tap, swipe, pinch), locators and locator composition, element state queries, assertions, screenshots/PDF, JavaScript evaluation, waiting mechanisms, viewport/emulation, network routing, storage (cookies, session), events, dialogs, file choosers, downloads, browser management, and debugging (trace, console, errors)
Unified code semantics tool (HIP-0300): Handles semantic code operations including parse (source to AST), serialize (AST to text with round-trip), symbols (list symbols in file), outline (symbols with imports/exports), definition (go to definition), references (find all references), search_symbol (find symbols across project), transform (codemod to patch), summarize (compress to summary), metrics (count files/lines by extension), exports (extract public exports), types (find type definitions), hierarchy (build class inheritance tree), rename (rename symbols across files), grep_replace (pattern replacement across files)
Descriptions violate length guideline (10-1024 chars): 'agent' is 1544 chars, 'browser' is 1789 chars. Guideline states too long descriptions waste tokens and bury key details. Current descriptions read as feature dumps rather than action-focused prompts.
No output schemas documented. Guideline requires documenting return types so LLMs know what fields to expect for downstream tool calls. No visibility into what agent/browser/code/config tools actually return.
'agent' tool combines 10 distinct actions (run, dag, swarm, consensus, dispatch, list, status, config, zen, review) into one tool. Violates single-responsibility principle, should split into separate tools for each orchestration pattern or at minimum separate control-plane tools (list, status, config) from execution tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Git-style configuration tool (HIP-0300): Actions get (default), set, list, toggle over three planes: index.scope (indexing scope: project|global|auto in ~/.hanzo/mcp/index_config.json), tools.<name>.enabled / enabled_tools.<name> (tool execution flags in ~/.hanzo/settings.json), and legacy per-indexer enable flag in index_settings block. Key resolution and output strings match Python implementation for interchangeability.
No error handling guidance. Tools provide no recovery hints, retryability classification, or actionable error messages. Guideline requires 'Error responses must tell the LLM what to do next' and categorize as retryable/user-fixable/fatal.
'browser' action parameter accepts 90+ enum values but descriptions are generic. Each action (navigate, click, screenshot, etc.) should clarify what it does, when to use vs related actions, and what it returns.
'config' tool mixes three configuration planes (index.scope, tools.*.enabled, legacy per-indexer flags) into single tool interface. Compounds complexity, should clarify which plane each key maps to and document interdependencies.
No security patterns visible: no audit trail declarations, no permission gate hints, no scope declarations (read/write/admin), no guidance on credential handling. Critical for production tools that may perform destructive operations (agent dispatch, browser automation).
Parameter 'timeout' appears in agent and browser without unit clarity or bounds. Agent timeout description says 'Timeout in seconds' but browser says 'milliseconds'. No min/max specified. Guideline requires 'Specify minimum and maximum for numeric parameters'.