A semantic routing MCP server that provides dog-related functionality (barking) and MCP server registry services. Includes a DoodleServer for dog-related queries and a RegistryServer for publishing and discovering MCP servers.
This server has 6 tools with highly uneven quality. Tools 1-3 (bark, get_last_bark, get_description) are demo/trivial with minimal descriptions. Tools 4-6 (register_server, heartbeat, discover_servers) have better structure but lack actionable descriptions. Only tools 4-6 have meaningful input schemas visible in code. Critical issues: (1) No tool descriptions explain WHEN to use each tool or what problem it solves. Descriptions are either missing or generic one-liners. (2) Parameter descriptions are sparse, the 'intensity' enum in bark lacks guidance on what intensity values mean in a dog-barking context. (3) No output schemas documented anywhere. What does bark() return? The code shows {sound, timestamp}, but agents have no schema to parse. (4) No error handling guidance, no tool tells the agent what to do if it fails. (5) Tool names are inconsistent in style: 'bark' vs 'register_server' vs 'get_last_bark'. (6) The registry tools (4-6) mix concerns: register_server also handles metadata and endpoint validation, which could fail silently. Source: doodle_server.py shows async methods but no MCP-layer schema definitions visible; mcp_registry/server.py defines input parameters in function signatures (not JSON Schema), making schemas implicit and unverifiable.
Simulate a dog bark with different intensities.
Discover available MCP servers, optionally filtered by capability
Get the semantic routing description for this server
Get information about the last bark
Update server heartbeat timestamp
Register a new MCP server with the registry
Missing output schemas for all tools. Code shows methods return dicts (e.g., {'sound': bark, 'timestamp': timestamp}) but no JSON Schema documents what fields agents should expect. This forces agents to parse unstructured responses and prevents downstream tool chaining.
Tool descriptions are too generic and do not explain WHEN to call them or HOW they differ from similar tools. E.g., 'Get information about the last bark' (get_last_bark) does not say: is this to retrieve timing, to verify a bark occurred, or to monitor state? No context for agent selection.
Parameter 'intensity' in bark() has no description explaining what 'quiet', 'normal', 'loud' mean in context or what the agent should prefer. Default is 'normal', is that safe? Will it startle a user? No guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 32 | - | v1 |
No error handling guidance in any tool. If register_server fails because server_id already exists, the code raises RegistryError with a message, but the description does not tell agents 'try discover_servers() to check if it exists' or 'use a different server_id'. Agents have no recovery path.
Tools 4-6 accept complex parameters (e.g., 'capabilities': List[str], 'metadata': Dict) but provide no constraint guidance. How many capabilities? What keys/values are valid for metadata? Are they validated server-side? No constraints in descriptions.
Inconsistent naming convention. Some tools use verb_noun (register_server, discover_servers), others use verb_noun with object (get_last_bark, get_description). Pattern is irregular, LLMs will spend reasoning cycles disambiguating similar-sounding names.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible. register_server and heartbeat are WRITE operations but tool definitions do not declare this, agents cannot distinguish safe retries from destructive calls.