Multi-module MCP server for customer support operations including policy document processing, defect analysis, database verification, and knowledge graph management
This server has 15 tools with mixed definition quality. Most tools have descriptions (11/15 present and substantive), but several critical gaps exist: (1) parameter descriptions are inconsistent, many parameters lack clear guidance on format, constraints, or purpose; (2) output schemas are completely undocumented, no tool declares what fields it returns or their types; (3) naming is generally clear (verb_noun pattern mostly followed), but some tools like 'find_order_by_invoice_number' and 'find_order_by_order_invoice_id' are confusingly similar; (4) error handling is invisible, no indication of what errors these tools can raise or how to recover; (5) no evidence of input validation or constraint enforcement in descriptions. The server sits at the boundary between D and C: definitions exist, but lack the rigor and LLM-optimization needed for reliable agent use. Several tools (especially database queries) expose low-level parameters like 'order_invoice_id' without explaining what they are or how agents should obtain them.
Analyzes a product defect image using Gemini 3 Vision and returns a one-line description.
Test the Neo4j database connection.
Create or merge a node in the knowledge graph.
Execute multiple Cypher statements in sequence. Ideal for bulk graph construction from extracted policy rules.
Execute a read-only Cypher query against the knowledge graph.
Execute a write Cypher query (CREATE, MERGE, DELETE, SET, etc.).
Output schemas completely undocumented. No tool declares what fields it returns, field types, or data structures. LLMs cannot determine what data they have after a call, forcing blind reasoning and likely downstream failures.
Parameter descriptions lack actionable constraints. E.g., 'limit' parameters mention defaults and bounds in some tools but not others; 'verification_email' is optional but no guidance on consequences of omission; 'properties' and 'parameters' are described as JSON strings with no schema or examples of valid structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Find a single order by invoice number (exact match). Returns full hierarchy: Customer, Order Details, and Order Items.
Find a single order by order_invoice_id (exact match). Returns full hierarchy: Customer, Order Details, and Order Items.
Retrieve the current graph schema (node labels, relationship types, property keys).
Get detailed statistics about the knowledge graph.
List detailed line items for a specific order_invoice_id. This tool fetches all products/items associated with an order. It is useful when you have the Order ID but need to see exactly what was purchased (SKUs, current quantities, etc.).
List historical orders for a specific customer email. This tool searches the database for all orders associated with the provided email address. It performs a case-insensitive search.
Parses ALL PDF policy documents in a directory and combines them into a single Markdown file. Each page is prefixed with a marker for traceability: <!-- PAGE:filename:page:start_line:end_line --> Also generates a combined_policy_index.json for citation lookup.
Parses a single policy PDF using LlamaParse and returns hierarchical Markdown.
Decodes a base64 PDF, parses the text, saves the text to a file, and returns the content.
Similar tool names create ambiguity. 'find_order_by_invoice_number' vs 'find_order_by_order_invoice_id' are difficult to distinguish, the description hints at subtle differences (one is for 'orders.order_invoice_id', the other is for 'invoice_number'), but the names do not signal which is which. LLMs will struggle to pick the right one.
No error handling guidance. Tools provide no information about what errors can occur, when they are retryable, or what the LLM should do next. E.g., 'execute_cypher_query' gives no hint about malformed Cypher, database connection failures, or permission errors.
Dangerous tools lack confirmation/dry-run pattern. 'execute_cypher_write', 'execute_cypher_batch', 'delete'-style operations should support a dry-run mode or require explicit confirmation to prevent accidental data loss. No evidence of this in descriptions or schema.
Parameter type and constraint inconsistency. 'queries' in 'execute_cypher_batch' is described as 'JSON array of Cypher query strings' but has type 'string', the parameter is a string representation of JSON, not a native array. This forces the LLM to construct JSON syntax, which is error-prone. Native array type is better.
Opaque identifier requirements. Tools like 'find_order_by_order_invoice_id' require an 'order_invoice_id' but give no guidance on how the agent obtains this ID. Must a user provide it? Is it the same as an 'order_id'? Does 'list_orders_by_customer_email' return these IDs so they can be chained? Undocumented dependencies.
Cypher query tools require deep domain expertise. 'execute_cypher_query' and 'execute_cypher_write' expect agents to formulate valid Cypher syntax. No description of the graph schema, node labels, relationship types, or examples. Agents will struggle to construct correct queries without additional context (e.g., call 'get_graph_schema' first, but descriptions do not say so).
File path parameters accept both absolute and relative paths with no clarity on working directory or resolution rules. E.g., 'analyze_defect_image' accepts 'image_path' like 'C:/images/defect.jpg', what if the file does not exist? Is the path relative to the server's working directory or the client's? Ambiguity invites errors.