An Elixir/Phoenix-based MCP server that provides policy-driven command classification, approval workflows, and audit logging for shell commands and MCP tool calls
Dingleberry is an HTTP-based policy classification and approval system with 5 tools. Tools have names starting with action verbs (classify_, approve_, reject_, record_), descriptions are present and reasonably specific, and schemas are defined using Jido.Action with type annotations. However, there are significant gaps: (1) parameter descriptions are minimal or missing in some tools; (2) output schemas are documented in code (e.g., classify_tool_call shows output_schema with risk, rule_name, description) but not exposed in the HTTP API response, the serialize_tool() function only returns parameters_schema, not output schema; (3) no enum constraints on multi-value fields (e.g., 'risk' accepts :safe, :warn, :block but is typed as :atom without enumeration); (4) error handling returns generic messages with no recovery guidance; (5) no documentation of tool composition or chaining; (6) write operations (approve_request, reject_request, record_audit) have no dry-run or confirmation patterns. The codebase is well-structured with Jido.Action modules and proper Elixir idioms, but the HTTP surface layer is minimal.
Approves a pending command in the approval queue
Classifies a shell command against YAML policy rules
Classifies an MCP tool call against YAML policy rules
Records a command interception event to the audit log
Rejects a pending command in the approval queue
Output schemas are defined in Jido.Action modules but not exposed via the HTTP API. The serialize_tool() function in tools_controller.ex only returns name, description, and parameters_schema, it omits output_schema. LLMs cannot see what fields to expect from tool results, breaking downstream tool composition and forcing agents to guess at response structure.
Enum constraints on multi-value parameters are absent from schema definitions. The 'scope' parameter in classify_command accepts :shell, :mcp, :all but is typed as :atom with no enum constraint. The 'risk' and 'decision' parameters in record_audit are free-form strings with no enumeration. LLMs cannot discover valid options and may hallucinate invalid values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 72 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions are sparse or absent. While parameter names exist, many lack actionable descriptions. Example: 'scope' in classify_command has description 'Classification scope (:shell, :mcp, :all)' but does not explain WHEN to use each scope or what the difference is. 'decided_by' in approve_request defaults to 'human' but lacks description of what values are accepted or why the field exists.
No error handling or recovery guidance. The tools_controller.ex run() action returns generic error messages ('ok: false, error: ...') with no instruction for the LLM on how to recover. Missing: pattern:recovery-guide and pattern:error-classification. An LLM hitting an invalid request_id receives an error with no guidance on whether to retry, lookup a valid ID, or escalate.
Write operations lack dry-run, confirmation, or idempotency declarations. approve_request, reject_request, and record_audit are state-modifying tools with no confirmation step or dry-run mode. If an LLM makes a mistake calling approve_request, the action is irreversible. No documentation of whether these operations are idempotent (can the same request be approved twice?) or what the expected state machine is.
Tool composition and chaining are not documented. There is no guidance on the intended workflow: e.g., call classify_command first, then (based on risk) call approve_request or reject_request, then record_audit. The HTTP API exposes tools individually with no dependency or sequencing hints. Per pattern:tool-chain, tool A's output must contain IDs that tool B needs, unclear if this holds (does classify_command return request_id that approve_request expects?).
Ambiguous tool distinctions. classify_command and classify_tool_call both perform classification against YAML policy rules but operate on different input types. Their names are similar and descriptions use the same phrasing ('Classifies ... against YAML policy rules'). Per naming guidance, similar-purpose tools must make distinctions obvious. The descriptions should clarify: when to call classify_command vs classify_tool_call?