A distributed multi-agent system that orchestrates complex tasks across specialized AI agents using A2A (Agent-to-Agent) protocol, MCP (Model Context Protocol) and LangGraph.
This MCP server exposes 7 basic math tools with consistent but minimal documentation. All tools have descriptions and visible input schemas, but descriptions are generic (mostly 10 - 30 words) and lack actionable guidance for LLM-driven selection. Parameter descriptions exist but are extremely terse (1 - 5 words each). No output schemas are documented. Error handling is absent, tools will crash on division by zero with no recovery guidance. Naming is reasonable (verb_noun) but lacks disambiguation cues. The server uses HTTP transport via FastMCP/Uvicorn, which is current. However, no tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite the tools being READ_ONLY. No pagination, no batching, no advanced composition patterns. Overall: a mediocre implementation that functions but lacks production-grade quality.
Add two numbers together. Example: add(3, 5) → 8
Calculate the cube of a number (number multiplied by itself twice). Example: cube(3) → 27
Divide the first number by the second number. Returns a floating-point result. Example: divide(20, 4) → 5.0
Multiply two numbers together. Example: multiply(7, 6) → 42
Raise the first number to the power of the second number (exponentiation). Example: power(2, 5) → 32
Calculate the square of a number (number multiplied by itself). Example: square(9) → 81
Subtract the second number from the first number. Example: subtract(10, 4) → 6
Descriptions are under 50 characters and lack actionable guidance. E.g. 'Add two numbers together. Example: add(3, 5) → 8' (48 chars) tells the LLM WHAT but not WHEN or WHY to call it. Baseline for A+ tools is 194 chars (p10=34, p90=392). These descriptions fall in the lower percentile and provide no differentiation from similar tools.
Parameter descriptions are minimal (1 - 5 words). E.g. 'First number' for parameter 'a' in add(). Baseline is 72 chars for A+ tools. LLMs cannot distinguish 'a' and 'b' reliably without richer guidance. The description should include type hints, constraints, and examples of valid values.
No output schemas are documented. Tools return integers or floats, but the response structure is not declared. LLMs need to know what fields to expect and their types. All 7 tools lack this.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
No error handling or recovery guidance. divide() will crash on division by zero with no actionable error message. Pattern: recovery-guide requires 'If X, then call Y' or 'Try this alternative.' None of these tools offer that.
No tool annotations despite tools being READ_ONLY. The spec defines readOnlyHint, destructiveHint, and idempotentHint annotations. All 7 math tools should be annotated as readOnlyHint=true to signal they are safe to call multiple times without side effects.
Parameter naming lacks type suffixes. Parameters are named 'a' and 'b', generic and ambiguous. Baseline rubric recommends 'base', 'exponent', 'dividend', 'divisor' etc. to self-document intent. For tools like power(a, b), unclear which is base and which is exponent without reading the description.
No input validation or constraint documentation. integer types are declared but no min/max bounds. divide() accepts any integer, including 0, which will crash. Baseline: numeric parameters should specify ranges (e.g. 'divisor must be non-zero').