Universal MCP Gateway for aggregating MCP servers into one HTTP/SSE endpoint with token optimization and a web dashboard
MCP Gateway exhibits mixed quality across its 9 tools. Tool naming follows verb_noun conventions (searchTools, getToolSchema, executeCode, callTool, createSkill, listSkills, executeSkill, deleteSkill), which is strong. However, descriptions vary significantly in quality and completeness. Most tools have basic descriptions (50-150 chars), but several lack crucial context about state modification, error handling, and when to use each tool. Input schemas are present and well-typed with enums and constraints (e.g., detailLevel in searchTools, mode in getToolSchema), but output schemas are undocumented, a critical gap for agent planning. Parameter descriptions are present but inconsistent in depth. The executeCode tool (marked IRREVERSIBLE) lacks explicit confirmation/dry-run patterns despite its destructive capability. Error handling guidance is absent across all tools. Security concerns: executeCode has no visible sandboxing validation in the schema definition itself, though source hints at VM isolation.
Call a tool from any connected backend MCP server
Save a reusable skill (code pattern) for future use
Delete a saved skill
Execute TypeScript/JavaScript code with access to MCP tools in a sandboxed environment
Execute a saved skill with optional parameters
Get full schema for a specific tool with lazy loading capability
Get tools organized in a hierarchical tree structure grouped by backend server
List all saved skills
Output schemas are completely undocumented. Tools return results but no schema describes what fields LLMs should expect, forcing agents to parse responses blindly and preventing downstream tool composition.
executeCode and executeSkill (IRREVERSIBLE operations) lack confirmation or dry-run patterns. No mechanism prevents accidental code execution or skill invocation. An agent in a retry loop could execute the same code multiple times with unintended side effects.
Error handling guidance is absent. Tool descriptions do not explain what errors might occur, how to recover, or what the agent should try next. A failed executeCode call returns no guidance on retry vs escalation.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | 2024-11-05+ | v1 |
Search and filter tools with various filtering options including query, backend, prefix, category, and detail level
Parameter descriptions lack depth. 'Parameters to pass to the skill' (executeSkill) tells the agent nothing about expected structure, types, or constraints. Compare to baseline: 72+ chars per param description is standard.
No tool chaining metadata. If searchTools returns tool results, downstream tools (getToolSchema, callTool) need tool IDs/names from the results. No documentation shows what fields bind searches to calls.
Pagination support is minimal. searchTools offers limit/offset but does not document a total count or hasMore flag in the response schema, forcing agents to guess if more results exist.