MCP server for customer success operations, providing tools for knowledge base search, ticket creation, customer history retrieval, escalation management, and multi-channel message delivery (email, WhatsApp, web)
Server demonstrates basic tool structure with 5 tools covering customer success workflows (search, ticket creation, history retrieval, escalation, response sending). However, significant quality gaps exist: descriptions are present but generic (mostly 50-90 chars, below the 194-char baseline); parameter descriptions are sparse or missing entirely; output schemas are undocumented; no error handling guidance; no security controls visible. The code shows in-memory fallback storage and Pydantic input models, indicating some schema awareness, but the MCP server registration and tool definition in src/agent/mcp_server.py (referenced but not shown) cannot be fully verified. Based on visible code fragments, tool definitions appear to be inferred rather than explicitly registered, triggering the per-tool cap at 50.
Create a support ticket for a customer
Escalate a conversation to human support
Retrieve customer history and ticket information
Search the product knowledge base for relevant information
Send formatted response to customer via specified channel
Tool descriptions are generic and underdeveloped. 'Search the product knowledge base for relevant information' (73 chars) and 'Create a support ticket for a customer' (37 chars) lack actionable context on WHEN to use vs similar tools, WHAT data is returned, and error handling paths. Baseline is 194 chars; these are 20-40% of expected length.
Parameter descriptions are missing or minimal. 'channel' param across multiple tools has description 'Communication channel (email, whatsapp, web)' but lacks guidance on what happens if an unsupported channel is passed. 'top_k' in search_knowledge_base is described but lacks constraint explanation (why max 10? is it a hard limit or guidance?).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Output schemas are not documented in visible tool definitions. The code shows tools return JSON strings (e.g., create_ticket returns '{"ticket_id": ..., "status": ..., "success": ...}') but no schema documentation is visible. LLMs cannot plan downstream calls or parse responses reliably without knowing expected output structure.
No error handling guidance provided. Functions catch exceptions and log errors (e.g., 'Knowledge base search unavailable. Please contact support.') but do not classify errors as retryable, user-fixable, or fatal. An LLM cannot decide whether to retry, ask the user, or proceed differently without this classification.
No validation error messages provide actionable feedback. If 'priority' is passed as 'P4' (invalid), no tool shows a message like 'Invalid priority: must be one of P1, P2, P3.' LLMs cannot self-correct without explicit constraint violations.
Tool definitions cannot be fully verified from visible code. The rubric states: 'If you cannot see the actual tool definition in the source (only inferred), cap that tool's overall at 50.' The actual MCP server registration in src/agent/mcp_server.py is not provided, only fragments.
No permission gates or security controls visible. Destructive tools (create_ticket, escalate_to_human, send_response) do not check whether the calling agent has authority. In-memory stores are unprotected global dicts.
Requires both customer_id AND email in get_customer_history but neither is required (both Optional). If both are omitted, function returns 'New customer' without error. Undocumented logic invites misuse.
Metadata field in send_response is an untyped dict with no schema. 'Additional message metadata' is vague, what fields are valid? What happens if invalid keys are passed?