Model Context Protocol (MCP) server for AgentHelm — AI Agent Control Plane. Provides tools for context retrieval, knowledge proposals, task management, incident tracking, and project brain operations.
AgentHelm MCP Server has moderate tool definitions with reasonable descriptions and schemas for most tools, but suffers from several quality gaps: inconsistent parameter documentation, missing output schema specifications, lack of error guidance, and weak error handling. Of 7 tools examined, all have descriptions and most have input schemas, but parameter quality is uneven. None of the tools document their output structures, which violates the 'Document the output schema' requirement. Tool descriptions are adequate (100-200 chars) but lack clear 'when to use' guidance and recovery hints. Several critical parameters lack descriptions (e.g., get_history's version/version_a/version_b fields are enumerated but parameter descriptions are minimal). No tool explicitly handles errors with actionable recovery guidance. The tools follow verb-noun naming conventions well, but lack error classification, confirmation patterns for destructive operations (propose_knowledge, record_incident), and per-tool output specifications.
Retrieves relevant, ranked context from the Project Brain using semantic context selection. Useful when starting a task, resolving design questions, or looking up project standards.
Retrieves version history, diffs, blame, or show details for the Project Brain. Allows reviewing exactly who proposed what, what files changed, what conflicts occurred, and when.
Retrieves a previously recorded incident by title or tags. Use this to check if a similar issue has been debugged before starting work.
Lists available tasks/checkpoints for the current project. Returns tasks with their latest checkpoint info so a fresh session can pick up where it left off.
Submits a Knowledge Proposal containing newly discovered or updated project design, decisions, schema changes, or API specifications. The proposal enters a validation queue before being compiled into the Project Brain.
No output schemas documented for any tool. LLMs cannot determine what fields to expect from responses, forcing them to guess field names and types. This breaks downstream tool chaining and wastes tokens on discovery.
get_history tool has required parameters (version, version_a, version_b) with minimal or missing descriptions. The 'action' enum is documented, but version parameters lack format, type, or range information. LLMs cannot infer whether these are sequential integers, UUIDs, or timestamps.
list_tasks has empty input schema (no properties). The tool description is vague ('Lists available tasks/checkpoints for the current project'). No indication of pagination, filtering, or result structure. Agents cannot determine if they can page through results or filter by date/status.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2026-07-28+ | v2 |
Records an incident/debugging finding to prevent duplicate work. Stores root cause, solution, and context so future sessions (or other agents) can retrieve it instead of re-debugging.
Resumes a previous agent session by retrieving the last completed checkpoint for a given task_id. Returns the state snapshot, step name, and output data to continue where the previous session left off.
propose_knowledge accepts arrays (apis_affected, db_changes, architecture) with type 'object' but no itemDescription. LLMs cannot determine what fields belong in each object. Should use additionalProperties schema or define a schema for each array item.
No error handling or recovery guidance in any tool. If propose_knowledge fails validation, get_history returns no matching version, or record_incident encounters a conflict, no guidance is provided on what to do next. Violates 'Error responses must tell the LLM what to do next' rule.
Destructive or high-impact tools (propose_knowledge submits to a validation queue; record_incident permanently stores data) lack confirmation patterns or dry-run modes. Agents could accidentally propose incorrect knowledge or mis-record incidents without a review step.
get_incident accepts 'title' and 'tags' as search criteria but provides no guidance on exact vs. partial matching, case sensitivity, or tag matching logic (AND vs OR). Description does not clarify format or constraints.
get_context requires 'task_hint' string but lacks guidance on length, format, or examples. Should specify min/max length and example hints (e.g., 'database migration', 'auth cookies') to help LLMs formulate effective queries.
No pagination parameters in list_tasks despite it being a discovery tool. If 100+ tasks exist, LLMs cannot determine how to retrieve all of them. Should support limit/offset or cursor-based pagination per the paginated-result pattern.