Enterprise Agent Supervisor - Rules Engine MCP Server for AI Agent Governance
ManagerProtocol has 18 tools with partial schema coverage and mixed description quality. Strengths: all tools have descriptions (avoiding 0-score floor); most have named parameters with types. Weaknesses: (1) many parameter descriptions are minimal or missing, causing schema scores to drop; (2) output schemas are not documented anywhere in the visible code, only input schemas are shown; (3) error handling and recovery guidance is absent; (4) several tools have vague or generic descriptions that don't answer WHEN or WHY to use them; (5) no tool annotations (readOnlyHint, destructiveHint) despite clear read-vs-write semantics marked in the source; (6) the server accepts secrets/tokens via environment but doesn't gate access or document permission requirements. Average per-tool score is 57.8, placing this in the 'Fair/Poor' range (C-/D+). This is typical for governance-heavy internal tools that lack agent-facing polish.
Add or update a rate limit configuration. Controls maximum requests per time window for specific action categories or globally.
Register a new custom business rule or override an existing rule. Validates rule structure before adding.
Apply all active business rules to a business context. Returns matched rules, aggregate risk score, and recommended constraints.
Approve a pending human approval request. Records approver and optional comments.
Create a GitHub issue using the gh CLI. Requires gh auth login. Can set labels, assignees, and body.
Evaluate CSS code against design system rules before adding to codebase. Returns violations, suggestions, and warnings.
Output schemas not documented. Tools define input schemas clearly (via Zod and JSON Schema), but nowhere in the source code are output schemas documented. LLMs cannot plan downstream calls or extract return fields without documented output structure. This violates pattern:tool and pattern:response-shaper.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The source code marks tools as 'READ_ONLY' or 'WRITE' in comments, but does not emit these as MCP tool annotations. Modern spec (2026-07-28) uses destructiveHint and readOnlyHint in tool registration to guide agent planning. This is a correctness issue for stateful systems.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
Deny a pending human approval request. Records denier and optional reason.
Evaluate an agent action against governance rules. Returns risk score, approval status, violations, and recommendations. Use this BEFORE executing any significant agent action to ensure compliance and safety. Returns: - status: approved | denied | pending_approval | rate_limited | requires_review - riskScore: 0-100 numeric score - riskLevel: critical | high | medium | low | minimal - violations: Array of rule violations - warnings: Array of warnings - requiresHumanApproval: Whether human approval is needed
Export audit log to JSON or CSV format with optional filtering. Downloads all matching events.
Retrieve audit events with filtering by type, agent, session, user, outcome, and date range. Supports pagination.
Get audit statistics including total events, breakdown by outcome, and time series data.
Verify supervisor is running and all systems are operational. Returns status of rules engine, audit logger, rate limiter, and database.
List all active governance rules with filtering and pagination. Returns rule details, counts, and metadata.
Load a preset rule configuration (minimal, standard, strict, financial, healthcare, development, frontend). Replaces current rules.
Log an audit event to the audit trail. Captures action taken, outcome, and metadata for compliance and debugging.
Disable or remove a governance rule by ID.
Request human approval for an action. Creates an approval request that must be approved or denied by a human operator.
Update supervisor configuration including strict mode, risk thresholds, approval requirements, and feature toggles.
Minimal parameter descriptions for governance rules. Parameters like 'conditions' and 'actions' in add_rule are typed as arrays but lack descriptions explaining what structures they accept, valid condition operators, or action types. LLMs cannot construct valid rule payloads without examples or explicit format documentation.
No error handling guidance in descriptions. Every tool description omits what errors can occur, how to recover, and whether calls are retryable. E.g., 'require_human_approval' could time out waiting for user input or fail if approval requests are disabled, but the description doesn't warn the agent.
Generic descriptions that don't answer WHEN or WHY. Example: 'apply_business_rules' says 'Apply all active business rules to a business context.' This is circular, it doesn't explain when the agent should call it instead of evaluate_action, what output it produces, or what 'matched rules' actually means in the agent's plan.
No permission or scope declaration. Tools like 'remove_rule', 'update_config', and 'add_rule' modify governance state but don't declare what permissions they require (e.g., 'admin:rules', 'write:config'). This violates pattern:scope-declaration and makes least-privilege configurations impossible.
Environment variables used for secrets/config but no validation. The code reads AUDIT_DB_PATH and NODE_ENV from environment, and instantiates supervisor with these values without validation. If AUDIT_DB_PATH is maliciously set to a path traversal like '../../../etc/passwd', the server could write audit logs to unintended locations.
Truncation warnings mentioned but not clear in tool responses. The code includes a limitResults() helper with truncation warnings and pagination hints, but it's unclear whether all list tools (list_rules, get_audit_events, etc.) actually use it. If some tools return unbounded results, context can be exhausted.
Inconsistent parameter naming. Some tools use underscores (e.g., 'agentId', 'userId'), others use camelCase (e.g., 'windowMs'). This inconsistency forces LLMs to reason about naming conventions and increases the likelihood of passing wrong parameter names.
No dry-run or confirmation for destructive operations. Tools like 'remove_rule', 'load_preset' (replaces current rules), and 'deny_request' are irreversible but lack a confirmation or dry-run step. An agent in a retry loop could accidentally wipe governance rules.