This is a demo/observability toolkit server, not a production-ready MCP server. Most tools are stub implementations returning mock data. Schemas are present but descriptions are minimal (10-50 chars, well below the 194 char baseline). Tool naming is reasonable but descriptions lack context about WHEN to use each tool and what dependencies exist. No output schema documentation visible. Error handling is absent, tools return mock success responses. 9 of 13 tools are located in demo files (real_world_server.py, server.py) which are not production code paths. The two tools with slightly better structure (transfer_funds_propose/commit) show awareness of proposal-commit patterns but lack full implementation detail. Overall: this reads as a reference implementation or testing framework, not a tool server meant for production agents.
Cancel a shipment with specified reason
Change subscription plan for an account with effective date
Create an expedited shipment for an order using specified carrier
Execute a transfer with specified amount and destination
Freeze a payment card for a customer due to fraud or other security reasons
Initiate a wire transfer with specified amount, destination IBAN, and reason
Issue a refund for an invoice with specified invoice ID, amount, and currency
Descriptions are critically short (10-50 chars vs 194 char baseline). 'Initiate a wire transfer with specified amount, destination IBAN, and reason' lacks guidance on prerequisites, error scenarios, or when to use vs similar tools. LLMs cannot determine from these minimal descriptions whether to call initiate_wire_transfer or transfer_funds_propose for the same intent.
No output schema documentation. Tools return dict[str, Any] but LLMs do not know what fields to expect. For example, initiate_wire_transfer returns {'status', 'operation', 'amount', 'destination_iban', 'reason'} but this is not documented anywhere. Without output schema, agents cannot plan downstream tool calls or extract the data they need.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Publish a feature flag with specified rollout percentage
Reset enterprise password for an employee with temporary secret
Schedule a clinic visit for a patient at specified time slot
Commit tool: executes side effects only with valid authorization artifact. Commits a previously proposed fund transfer.
Proposal tool: never executes side effects. Creates a proposal for fund transfer with amount and recipient.
Unfreeze a payment card for a customer after approval via ticket
No error handling guidance. All tools are mocks that return success. Real errors (invalid IBAN, insufficient funds, authentication failure, permission denied) are not documented. LLMs do not know what error codes to expect or how to recover. This violates the recovery-guide pattern.
Parameter descriptions are missing or generic. 'amount' is described only as 'Amount to transfer' with no currency, range, or precision constraints. 'destination_iban' has no format specification or validation hint. 'rollout_percent' lacks constraint that it must be 0-100 (though the type hint int is present).
Irreversible operations (initiate_wire_transfer, reset_enterprise_password, publish_feature_flag) are not gated behind confirmation or dry-run steps. Per pattern:confirmation-request, irreversible tools should support a proposal+commit pattern or require explicit approval. This server provides transfer_funds_propose/commit as an example but does not apply the pattern to the high-risk wire_transfer tool.
No parameter validation or constraint documentation. 'slot_iso' for schedule_clinic_visit is described only as 'ISO 8601 time slot for visit' with no validation that it is a future date, within business hours, or in a valid timezone. LLMs will pass arbitrary ISO strings and receive silent failures or unexpected behavior.
Tools located in demo files (/demo/real_world_server.py, /demo/server.py, /examples/) are not production implementations. They are stub return-mocks. This is acceptable for a testing/reference toolkit but should not be scored as if they were real callable services.
No chaining IDs in responses. E.g., initiate_wire_transfer returns operation ID but likely needs a reference ID or tracking number that downstream tools (check_transfer_status, cancel_transfer) would accept. Responses do not include IDs needed for follow-up calls.
Overlapping tool responsibilities. Both 'initiate_wire_transfer' and 'execute_transfer' exist (tools 1 and 11). Both appear to perform transfers but with slightly different semantics ('initiate' vs 'execute'). LLMs will waste reasoning cycles deciding between them. Consolidate into one canonical tool or document the distinction clearly.