MCP server that bridges Claude Chat to Claude Code running on the user's machine for automated task execution, code generation, and repository management
Herald demonstrates solid definition quality with consistent naming conventions (verb_noun pattern), comprehensive parameter descriptions, and well-structured schemas. All 10 tools follow the verb-first naming pattern (list_, start_, check_, get_, cancel_, read_, push_). Most parameters include type definitions and descriptions. However, there are notable gaps: output schemas are not explicitly documented in the visible code, some parameter descriptions lack concrete constraints (e.g., ranges for numeric fields), and error handling guidance is minimal. The tool definitions are properly registered with mcp-go and visible in source, avoiding the 'inferred tool' cap. descriptions range from 34-392 chars, averaging ~150 chars, which is good but some are verbose. Parameter annotations are present but could be more prescriptive about valid ranges and formats.
Cancel a running or pending task.
Check the current status and progress of a running task. Supports long-polling with wait_seconds to reduce polling overhead. IMPORTANT: Always use wait_seconds=30 to long-poll efficiently. This holds the connection server-side and returns immediately when the task status changes, avoiding wasteful rapid polling.
Show Git diff of changes. Use task_id to diff a task's branch against current branch, or project to diff uncommitted changes.
Get logs and activity history.
Get the complete result of a finished task.
Push the current Claude Code session context to Herald for remote monitoring and continuation. Call this when the user wants to continue working from another device.
Output schemas are not explicitly documented. Tool handlers exist in internal/mcp/handlers but return types are not visible in the source provided. LLMs cannot determine what fields to expect from tool responses, risking incorrect downstream composition.
Numeric parameters lack explicit range constraints. 'timeout_minutes', 'wait_seconds', 'limit', 'output_lines', 'turns', 'line_start', 'line_end' have no min/max documented. LLMs may pass unbounded values (e.g., timeout_minutes=999999) causing resource exhaustion.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all configured projects with their Git status and description.
List tasks with optional filters.
Read a file from a configured project (path-safe).
Start a Claude Code task on a project. Returns immediately with a task ID. The task runs asynchronously — use check_task to monitor progress. Tasks typically take 1-10 minutes. Use check_task with wait_seconds=30 to long-poll efficiently instead of polling rapidly.
Error handling and recovery guidance is absent. Tools like 'start_task', 'cancel_task' modify state but descriptions do not indicate failure modes, retryability, or what the LLM should do if a task fails to start or cancel. No error classification (retryable vs user-fixable vs fatal).
Parameter interdependencies not documented. 'start_task' has 'session_id' (to resume) and 'git_branch' (to create), but it's unclear: can both be specified? What if they conflict? Does session_id override git_branch? LLMs cannot reason about valid parameter combinations.
No pagination or result limits documented for list tools. 'list_tasks' and 'list_projects' accept 'limit' but baseline expectations are unclear. Default limits should be enforced and documented to prevent context window exhaustion.
'read_file' has no description of path safety guarantees. Can it read outside the project root? Is path traversal (..) prevented? LLMs may attempt to read /etc/passwd or other sensitive files if path validation is not explicit.