A secure, declarative MCP runtime for orchestrating agents with HTTP bridges, credential management, trust tiers, and audit logging.
Heddle presents a moderately structured MCP server with 13 tools covering development automation, GPU/VRAM management, and operational intelligence. Tool definitions are visible in source code with explicit parameter schemas and descriptions, which is positive. However, there are significant gaps: (1) Most tool descriptions are functional but brief (average ~60 chars), lacking the depth needed for LLM tool selection heuristics. (2) Parameter descriptions exist but are often minimal, assuming agent familiarity with domain concepts. (3) Output schemas are largely undocumented, tools describe what they return in prose but lack formal schema declarations. (4) Error handling guidance is absent; tools do not indicate recovery paths or error categorization. (5) Two tools (run_tui, send_keys) execute interactive terminal operations with minimal safety guardrails documented. (6) Several tools accept free-form string parameters (e.g., 'flags' in build, 'pattern' in run_tests) without documented constraints, inviting invalid LLM input. The codebase shows thoughtful security infrastructure (trust enforcer, audit logging, credential broker, rate limiting, escalation engine) but these are not reflected in tool-level documentation for the agent's perspective.
Build a Go project. Returns exit code, stdout, and stderr.
Capture the current terminal contents of a tmux session.
Generate the daily operations briefing. Gathers data from all sources, then uses Ollama to synthesize.
Get git status for a project: branch, dirty files, last 5 commits.
List all models across Ollama and GGUF library.
Read a file from the filesystem. Path can use ~ for home directory.
Run Go tests. Returns full test output.
Output schemas are not formally documented. Tools like vram_status, list_all_models, daily_briefing, system_health_check, and threat_landscape have no visible return type schemas in the source. Tools return structured data but LLMs cannot reliably parse outputs without formal schema declarations.
Parameter constraints are underdocumented. The 'flags' parameter in build (accepting arbitrary go build flags) and 'pattern' parameter in run_tests lack documented format constraints, regex patterns, or examples. This invites invalid LLM input like 'flags: -race -v -invalid-flag'.
Error handling guidance is absent across all tools. Tool descriptions do not indicate error conditions, recovery strategies, or categorization (retryable vs. user-fixable vs. fatal). For example, if git_status fails because the project path does not exist, the agent has no guidance on what to do next.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 67 | <=2025-11-25 | v2 |
Spawn a TUI binary in a detached tmux session for interactive testing.
Send keystrokes to a running tmux session.
Intelligently load a model, evicting others if needed.
Quick health check — just Prometheus data.
Synthesized threat landscape from intel-rag.
Comprehensive VRAM and GPU status report.
Interactive terminal tools (run_tui, send_keys) lack confirmation/dry-run safeguards in their documented interface. These are destructive operations (spawning processes, sending keystrokes) but their descriptions do not indicate how to preview or confirm actions before execution.
Tool descriptions are too brief. Average length ~60 chars; production baseline is 194 chars. Examples: 'Spawn a TUI binary in a detached tmux session for interactive testing' (66 chars) lacks context on when to use this vs. alternatives, prerequisites (tmux must be running), or side effects.
Pagination and result limits are not documented for tools that return lists. list_all_models and threat_landscape (which synthesizes data from intel-rag) do not specify maximum result count, pagination mechanism, or how agents should handle large result sets.
Parameter relationships are not documented. For example, run_tui's 'session' and 'binary' parameters have a dependency: the session must exist (or be created) for send_keys to work on it later. This multi-step composition is not explained in parameter descriptions.