MCP server for Stripe payment processing integration
This server has 6 tools with explicit schemas and descriptions visible in src/tools.py. However, significant quality issues prevent a higher score: (1) parameter descriptions are largely missing or minimal, most parameters lack context for LLM selection; (2) output schemas are not documented anywhere; (3) descriptions are short but adequate (mostly 20-60 chars, above the 20-char floor but below production baseline of 194 chars); (4) no error handling guidance visible; (5) no tool annotations or risk metadata exposed via the MCP protocol; (6) field-naming inconsistencies (e.g., customer_id vs customer) may force LLM reasoning. Tools are well-named with clear verbs (customer_create, payment_intent_create, refund_create), and schemas include type information, but parameter descriptions are sparse and output structures are completely undocumented. This is typical of a functional but not production-ready MCP server.
List recent charges
Create a new customer in Stripe
Retrieve a customer's details
Update customer information
Create a payment intent for processing payments
Create a refund for a charge
Output schemas completely undocumented. Source code shows tool definitions but no response structures are defined. LLMs cannot plan follow-up calls or extract required fields (e.g., created customer_id for subsequent updates).
Parameter descriptions missing on most inputs. customer_id, name, metadata, currency, customer, limit, customer_id (charge_list), charge_id, amount (refund) lack any description. Tool descriptions are also minimal (avg 35 chars vs baseline 194 chars). LLMs cannot disambiguate intent.
update_fields parameter in customer_update accepts arbitrary object with no schema constraints. This is dangerous, LLMs may pass invalid field names, nested structures, or system-reserved fields without validation feedback.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
No pagination guidance or schema for charge_list. The tool accepts limit and customer_id but offers no total_count, next_cursor, or offset to support iterating large result sets. No documentation of list size cap or behavior when limit is exceeded.
Destructive operations (refund_create, customer_update, payment_intent_create) lack dry-run, confirmation, or explicit error recovery guidance. No error_classification pattern, LLMs cannot determine if a failure is retryable or fatal.
Tool descriptions do not explicitly state when each tool modifies state. 'Create a new customer', 'Update customer information', and 'Create a refund' are actions but lack clear guidance on irreversibility. Agents need explicit: this is a WRITE, this is REVERSIBLE, this is IRREVERSIBLE.
No natural identifier support. All tools require IDs (customer_id, charge_id). If an agent knows 'acme@example.com' but needs customer_id, it must call customer_retrieve by ... what? No search_customers tool exists. Agents cannot resolve names/emails to IDs.
No tool annotations visible in MCP definitions. Risk metadata (Risk: WRITE, REVERSIBLE) is present in the evaluation spec but not exposed via tool annotations (readOnlyHint, destructiveHint, idempotentHint) that MCP clients can parse.