A routing proxy for agent traffic to MCP servers that manages connections to multiple MCP servers and exposes their tools via HTTP and MCP protocols
MCP Gateway is a routing proxy that exposes 5 tools with significant quality gaps. Three tools (get_current_time, calculate, echo) are defined in demo_server.py with basic schemas and descriptions. Two tools (search_tools, execute_tool) are inferred from examples/example_client_mcp.py without visible registration code. Tool descriptions are present but generic (10-50 chars). Parameter descriptions lack depth and constraint documentation. No output schemas are documented. The calculate tool uses unsafe eval() despite claiming 'safe evaluation'. Error handling is minimal and does not guide recovery. The architecture mixes concerns: demo tools, client examples, and gateway routing in a single repo make tool intent unclear. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. The server is marked as alpha (v0.1.0) and shows signs of incomplete patterns.
Perform basic arithmetic calculations
Echo back the provided message
Execute a specific MCP tool on the appropriate server
Get the current date and time
Search for MCP tools across all connected servers
Two tools (search_tools, execute_tool) are inferred from example client code, not explicitly registered via a visible @app.list_tools() or equivalent. Tool definitions are not in the source code.
No output schemas documented for any of the 5 tools. LLMs cannot plan downstream operations or extract required fields (e.g., tool_id for chaining).
Tool 'calculate' uses eval() for expression evaluation despite claiming 'safe evaluation'. Character whitelist can be bypassed via encoding tricks or numeric literals that evaluate to code.
Parameter descriptions are vague or missing. Examples: 'Search query string' (13 chars), 'Message to echo back' (20 chars, borderline). No explanation of constraints, formats, or allowed values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Tool 'execute_tool' accepts 'args' as unconstrained object. No nested schema, no type checking, no examples of valid structures. Increases risk of malformed invocations and makes LLM selection ambiguous.
No error handling guidance in tool descriptions. If calculate() fails on invalid input, if search_tools() finds no matches, or if execute_tool() hits a downstream error, descriptions do not tell the LLM what to do next.
Tool 'execute_tool' is marked as WRITE (destructive) but description does not mention confirmation, dry-run, or permission checks. Agents may irreversibly delete data without warning.
Tool names 'search_tools' and 'execute_tool' are generic and could collide or confuse with local/remote variants. No clarification on scope (gateway-wide? session-specific? backend-specific?).
No mention of pagination, result limits, or total counts for search_tools(). If 100+ tools are available, returning all at once wastes tokens and crashes context.
Tools lack idempotency guarantees. Calling calculate('2+2') twice should return the same result, but get_current_time() returns different times on each call. If agents retry on transient errors, duplicate time-sensitive operations may occur.
Descriptions use example values ('e.g., 2+2, 10*5' for calculate; 'e.g., 03:45:30 PM' for time format). LLMs tend to reuse example values literally, causing copy-paste errors in real calls.