A comprehensive Model Context Protocol server built in Go with advanced features including AI routing, compliance tracking, feature flags, task management, and observability
MCP Ultra exposes 29 tools across task management, feature flags, AI routing, compliance tracking, and health checks. While tool names follow verb_noun conventions (CreateTask, ListTasks, UpdateTask, etc.), most tool descriptions are extremely brief (10-25 chars), far below the 50-200 char baseline for LLM-optimized descriptions. Input schemas are present but lack detailed parameter descriptions, many parameters (e.g., 'status', 'strategy', 'attributes') have no explanation of valid values, ranges, or constraints. Output schemas are entirely undocumented. Error handling guidance is absent. The server implements appropriate risk tagging (READ_ONLY, WRITE, DESTRUCTIVE) but does not expose this via tool annotations in the MCP protocol. Overall, the tools are structurally sound but severely under-documented for LLM consumption.
Mark a task as complete
Create a new feature flag
Create a new task
Make AI routing decision based on use case, selecting provider and model
Delete a feature flag
Delete a task
Automatically discover data sources and their mappings by scanning databases, APIs, and files
Evaluate a feature flag for a specific user
Tool descriptions are critically short (10-30 characters), far below the 50-200 char LLM-optimized baseline. Examples: 'Create a new task' (17 chars), 'List all tasks with optional filtering' (38 chars). LLMs cannot determine when/why to select these tools.
No parameter descriptions provided in any tool. Parameters like 'status', 'strategy', 'attributes', 'parameters', 'legalBasis', 'retention' are ambiguous without explanation. LLMs cannot infer whether 'status' is a Boolean, enum, or free-text field. This violates the critical rule that every parameter needs a description explaining what it controls.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Generate a comprehensive data map with all field mappings and inventory items
Retrieve all data field compliance mappings
Retrieve the compliance mapping for a specific data field
Retrieve a specific feature flag by key
Retrieve a specific task by ID
Retrieve tasks assigned to a specific user
Retrieve tasks filtered by status
Get overall health status of the service
List all feature flags
List all tasks with optional filtering
Kubernetes liveness probe endpoint
Map a data field with its compliance metadata including PII type, sensitivity, legal basis, and retention rules
Publish inference error event
Publish inference summary event with token counts, latency, and cost metrics
Publish policy block event when a policy rule is violated
Publish AI router decision event with provider and model selection reason
Kubernetes readiness probe endpoint
Record access to data for compliance tracking with actor, action, and purpose
Track how data flows through the system from source to destination
Update an existing feature flag
Update an existing task
No output schemas documented. LLMs cannot plan downstream tool calls or extract data without knowing response structure. Baseline expectation: 100% of A+ tools document return types. No evidence of documented return structures for any of the 29 tools.
Pagination support unclear. ListTasks and GetTasksByStatus accept 'limit' and 'offset', but descriptions do not specify valid ranges (e.g. 1 - 100), default values, or max results. Baseline: list tools should clearly document pagination contract and return total count. No evidence of these safeguards.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are not visible in the MCP tool definitions, despite internal risk tagging in the server (READ_ONLY, WRITE, DESTRUCTIVE). The protocol expects these to be declared for each tool so clients can apply safety policies. LLMs need explicit hints about which operations are reversible.
No error handling guidance. No evidence of error classification (retryable vs user-fixable vs fatal) or recovery hints. Baseline: error responses must tell the LLM what to do next. E.g. 'User not found. Try search_users() with a partial name.' None of the tools provide this.
Parameter constraints not visible. Tools accept 'strategy' (CreateFlag, UpdateFlag), 'severity' (PublishPolicyBlock), 'action' (RecordDataAccess) without enums or validation hints. LLMs will hallucinate invalid values. Baseline: declare known sets of values as enums and describe valid ranges in parameter descriptions.
Destructive operations (DeleteTask, DeleteFlag) lack confirmation/dry-run patterns. Baseline pattern: irreversible operations should support a confirm-before-execute step or explicit acknowledge parameter. No evidence of these safeguards.
Composite tools without clear decomposition: PublishInferenceSummary bundles token counts, latency, and cost. Consider splitting into separate tools or documenting why these must be published together. Baseline: each tool should do exactly one thing.
Compliance tools (MapDataField, TrackDataFlow, RecordDataAccess, DiscoverDataSources) accept complex nested objects (retention: object, source/destination: object, attributes: object) without documented schemas or examples. LLMs cannot construct valid payloads without clear guidance on structure.