An AI router that manages LLM protocols, MCP tools, and request routing with support for multiple LLM providers (Anthropic, AWS Bedrock) and MCP tool integration
Scoring was not performed
Duplicate tool functionality: Both 'adder' and 'add' perform identical addition operations with identical signatures. This violates the single-responsibility and composition patterns and forces LLMs to waste reasoning cycles choosing between them.
Descriptions lack actionable context. Tool descriptions are 40-80 characters and state only WHAT the tool does (e.g., 'Adds two numbers together'). They do not explain WHEN to use it, how it differs from similar tools, or what prerequisites exist. This is below the baseline 194-char average for production tools.
Missing output schema documentation. No tool documents its return type, structure, or fields. LLMs cannot infer what fields will be available for downstream tool calls or what data to extract from responses.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 64 | 2025-03-26+ | v1 |
No error handling guidance. Tools provide no descriptions of failure modes, error conditions, or recovery steps. The 'fail' tool has only a 4-character description and empty input schema, it is clearly test-only and offers no production value.
Filesystem tool lacks security context. The 'filesystem' tool accepts 'operation' enum with delete capability but provides no description of permission checks, path validation, or constraints. A destructive tool without explicit safeguards violates the permission-gate pattern.
Parameter descriptions are minimal. Most parameters have 1 - 2 word descriptions (e.g., 'First number to add', 'System environment variable name to retrieve'). Baseline for production tools is 72 chars average. No descriptions explain constraints, formats, or valid ranges.
Generic tool naming undermines clarity. The 'echo' tool has a clear purpose, but 'fail' is vague, does it test error paths, intentionally break something, or simulate a failure mode? Names like 'test_error_handling' or 'simulate_tool_failure' would be more self-documenting.